Metrics

041 · METRICS

Metrics

ML experiment tracking built for agents

Experiment tracking your agents can use.Log from Python in one line. Read it back as JSON, SQL or Python. Free with fair use.

Read the docsAgents: /llms.txt · curl any page
11 µs
per metric() call, Python on Linux
0.24 s
to read a series, median, across the Atlantic
11 s
to pull 118 runs into DuckDB, cold

01 · Agent first

If you can click it, your agent can type it.

Everything in the app is a command. syvain-metrics finds runs, reads and draws their curves, files and names them, and hands you a link to the exact charts it wants you to see. Every answer is JSON.

zsh
$ syvain-metrics experiment list --search lr3e-4
{"experimentId":"8c1f…","slug":"mamba-lr3e-4-seed7","folderPath":"/mamba/lr-sweep"}
{"experimentId":"a02d…","slug":"mamba-lr3e-4-seed8","folderPath":"/mamba/lr-sweep"}
$ syvain-metrics series render mamba-lr3e-4-seed7 mamba-lr3e-4-seed8 \
    --name loss --filter split=valid -o loss.png
$ syvain-metrics experiment rename mamba-lr3e-4-seed7 "Mamba, best lr"
$ syvain-metrics view link --folder /mamba/lr-sweep
{"url":"https://metrics.041.io/app?org_id=…&w=…"}

02 · The collector

Log anything. Wait for nothing.

TRAIN.PY
from syvain_metrics_collector import Collector

exp = Collector().experiment("mamba-lr3e-4-seed7", meta={"lr": 3e-4, "seed": 7})

with exp.run():
    for step, batch in enumerate(loader):
        exp.metric("loss", train_step(batch), step, {"split": "train"})
        if step % 5_000 == 0:
            exp.annotation("checkpoint", {"path": save_checkpoint(step)}, step=step)
01

Simple

One import, one line per number. The key comes from the environment, the run's lifecycle from a with block.

02

Flexible

Any JSON as config. Split any series by metadata. Annotate checkpoints, evals and anything else that isn't a number.

03

Fast

A Rust core batches and ships on its own thread. 11 µs per call; a million points queued and written in 11 s.

03 · Read it back

Four ways in. Your agent picks.

The same runs over HTTP, from the shell, in a notebook or in SQL. Every point at full resolution, on a series database we built for this →

HTTP API

JSON in, streamed JSON out. Anything that speaks HTTP can read it.

POST /api/v2/query/series

Command line

Curves as JSON lines, or as a PNG for agents that read images.

syvain-metrics series query

Python

Curves straight into a notebook or an analysis script.

exp.series("loss")

DuckDB

Every run as SQL tables, cached on disk as Parquet.

SELECT … FROM m.series

04 · The app

And when you want to look yourself.

Charts take the screen. Loss sits on top, found by itself; every diagnostic follows, drawn as you scroll. Drag runs between folders, set limits and log scales per chart, and save the view for the whole team.

Free, with fair use.

Log what your research needs. Need more, or something we don't do yet? info@041.io