Metrics
All pages
Docs · Start hereMarkdown

Why Metrics

Metrics is experiment tracking for research where agents write, launch and read most of the runs. Logging stays one line of Python, and everything that was logged can be read back as JSON, SQL or Python objects without a browser. The app is still there for the moment a person wants to look.

What changes when agents run the research

A person reading a dashboard skims a hundred curves and remembers which one looked odd. An agent can't skim a picture it never fetched, and it can't click through a UI built for a mouse. It needs the same numbers the charts show, in a form it can parse, filter and compare, reachable from a shell on the machine where it works.

Agents also run more experiments than people do, in parallel, and in long-running loops. Tracking has to stay out of the training step, never lose a point quietly, and make the run's configuration and outcome as easy to query as its curves.

What Metrics does

  • Logs from Python with one line per metric. The collector queues each value in about eleven microseconds and delivers it from a Rust core on its own thread, so training never waits on the network. One flush_or_raise() at the end fails the job if a single point was lost.
  • Answers every question as JSON. The command line addresses experiments by slug and folders by path, prints newline-delimited JSON and exits with status 1 on errors. It can also render a chart as a PNG for an agent that reads images.
  • Treats runs as a database. The DuckDB extension attaches folders, experiments, annotations and series as tables with a local Parquet cache. The Python API client reads the same data from analysis scripts.
  • Keeps every point. Series are stored at full resolution in a series database of our own. Nothing is sampled or averaged on the way in.
  • Puts the charts first in the app. The app opens on the loss, draws every other metric below it and saves the setup as a view the team can open.
  • Documents itself for agents. Every page of these docs is markdown at /docs/<page>.md, any page answers curl with markdown, and /llms.txt lists them all. See working as an agent.

What we left out

  • Artifacts and checkpoints. Store them where they already live and put the path in an annotation.
  • System monitoring. Log the GPU numbers you need as metrics; nothing is collected behind your back.
  • Hyperparameter search. Your agent or your sweep tool decides what to run; Metrics records what happened.
  • Media and tables. A metric is a number at a step. Annotations carry text and JSON for everything else.

Price

Free, with fair use. If you need more, or something we don't do yet, email info@041.io. See limits and fair use.