All pages
Working as an agent
Metrics is built to be driven by an agent. Log with the collector, read with the command line or DuckDB, and read these docs as markdown. Every read answers in JSON or SQL rows, every failure has an exit code, and nothing asks anything interactively once a key is in place. This page is the loop and the rules.
Credentials
An agent works with an organization API key in SYVAIN_METRICS_API_KEY, or with
the login saved by syvain-metrics auth login on that machine. The CLI, the
Python API client and DuckDB read the environment variable first and the saved
login second; the collector reads only the environment variable. An agent can
also set everything up itself: auth login --device prints a link for a person
to open, then org create and api-key create need no browser. See the
quickstart and organizations and keys.
The loop
syvain-metrics folder create /mamba/lr-sweep # one folder per question
python train.py --slug mamba-lr3e-4-seed7 --folder … # logs with the collector
syvain-metrics experiment get mamba-lr3e-4-seed7 # status: running, done or error
syvain-metrics experiment catalog mamba-lr3e-4-seed7 # what was logged, with metadata values
syvain-metrics series query mamba-lr3e-4-seed7 --name loss --filter split=valid
syvain-metrics series render mamba-lr3e-4-seed7 --name loss -o loss.png- Name runs so a glob finds them. The slug is the experiment's identity.
Put the question, the variable and the seed in it:
mamba-lr3e-4-seed7. Reopening a slug continues the same experiment. - Record the configuration as
meta. Everything the run varied, as JSON, so a query can join results to hyperparameters instead of parsing slugs. - Put each question in a folder. Folders are paths, the CLI creates missing parents, and DuckDB and the app select whole folders.
- Check status before reading results.
experiment getsaysrunning,doneorerror, with the error message and the time of the last event. A run that stopped logging withoutdoneis stillrunning. - Read the catalog before querying. It lists every metric name with the metadata keys and values it was logged with, so filters use values that exist.
- Compare in SQL. For more than a couple of runs, attach DuckDB and query the folder; the local cache makes the second query free. See DuckDB.
Reading what happened
| want | read | never |
|---|---|---|
| did it finish | experiment get: status, error |
the last value of a curve |
| what it logged | experiment catalog |
guess metric names |
| the numbers | series query, or m.series in DuckDB |
the PNG |
| a picture | series render -o chart.png |
a screenshot of the app |
| why it stopped | experiment annotations, error |
the training log |
When a person wants to look for themselves, give them a link rather than a
description of where to click.
syvain-metrics view link --folder /mamba/lr-sweep prints an app link with
those runs on the charts; --state sets the grouping, axes and filters, and
view create saves it for the team. See
views and links.
Rules that keep results honest
- End every job with
flush_or_raise(). It fails the job when any point was dropped, rejected or not delivered, so a run that finished without an exception has all its data. - One name per quantity.
losswithmetadata={"split": "valid"}, notvalid_loss. See what to log. - Compare like with like. Filter by the same metadata values in every run, align on step, and check that the runs reached the same step before comparing final values.
- Don't trust a curve from a failed run. Check status first.
- Leave notes. An annotation records why something happened, next to the step it happened at: a restart, a changed data mix, a bug found.
Reading these docs
Every page is plain markdown. Fetch any URL on this site with curl, or ask for
Accept: text/markdown, and the answer is markdown:
curl https://metrics.041.io/docs/cli
curl https://metrics.041.io/llms.txt # every page, one line each
curl https://metrics.041.io/llms-full.txt # every page in one file