# Working as an agent

Metrics is built to be driven by an agent. Log with the
[collector](https://metrics.041.io/docs/collector.md), read with the [command line](https://metrics.041.io/docs/cli.md) or
[DuckDB](https://metrics.041.io/docs/duckdb.md), and read these docs as markdown. Every read answers in JSON
or SQL rows, every failure has an exit code, and nothing asks anything
interactively once a key is in place. This page is the loop and the rules.

## Credentials

An agent works with an organization API key in `SYVAIN_METRICS_API_KEY`, or with
the login saved by `syvain-metrics auth login` on that machine. The CLI, the
Python API client and DuckDB read the environment variable first and the saved
login second; the collector reads only the environment variable. An agent can
also set everything up itself: `auth login --device` prints a link for a person
to open, then `org create` and `api-key create` need no browser. See the
[quickstart](https://metrics.041.io/docs/quickstart.md) and [organizations and keys](https://metrics.041.io/docs/organizations.md).

## The loop

```bash
syvain-metrics folder create /mamba/lr-sweep              # one folder per question
python train.py --slug mamba-lr3e-4-seed7 --folder …      # logs with the collector
syvain-metrics experiment get mamba-lr3e-4-seed7          # status: running, done or error
syvain-metrics experiment catalog mamba-lr3e-4-seed7      # what was logged, with metadata values
syvain-metrics series query mamba-lr3e-4-seed7 --name loss --filter split=valid
syvain-metrics series render mamba-lr3e-4-seed7 --name loss -o loss.png
```

1. **Name runs so a glob finds them.** The slug is the experiment's identity.
   Put the question, the variable and the seed in it: `mamba-lr3e-4-seed7`.
   Reopening a slug continues the same experiment.
2. **Record the configuration as `meta`.** Everything the run varied, as JSON,
   so a query can join results to hyperparameters instead of parsing slugs.
3. **Put each question in a folder.** Folders are paths, the CLI creates missing
   parents, and DuckDB and the app select whole folders.
4. **Check status before reading results.** `experiment get` says `running`,
   `done` or `error`, with the error message and the time of the last event. A
   run that stopped logging without `done` is still `running`.
5. **Read the catalog before querying.** It lists every metric name with the
   metadata keys and values it was logged with, so filters use values that
   exist.
6. **Compare in SQL.** For more than a couple of runs, attach DuckDB and query
   the folder; the local cache makes the second query free. See
   [DuckDB](https://metrics.041.io/docs/duckdb.md).

## Reading what happened

| want           | read                                    | never                     |
| -------------- | --------------------------------------- | ------------------------- |
| did it finish  | `experiment get`: `status`, `error`     | the last value of a curve |
| what it logged | `experiment catalog`                    | guess metric names        |
| the numbers    | `series query`, or `m.series` in DuckDB | the PNG                   |
| a picture      | `series render -o chart.png`            | a screenshot of the app   |
| why it stopped | `experiment annotations`, `error`       | the training log          |

When a person wants to look for themselves, give them a link rather than a
description of where to click.
`syvain-metrics view link --folder /mamba/lr-sweep` prints an app link with
those runs on the charts; `--state` sets the grouping, axes and filters, and
`view create` saves it for the team. See
[views and links](https://metrics.041.io/docs/app.md#views-and-links).

## Rules that keep results honest

- **End every job with `flush_or_raise()`.** It fails the job when any point was
  dropped, rejected or not delivered, so a run that finished without an
  exception has all its data.
- **One name per quantity.** `loss` with `metadata={"split": "valid"}`, not
  `valid_loss`. See [what to log](https://metrics.041.io/docs/logging-guide.md).
- **Compare like with like.** Filter by the same metadata values in every run,
  align on step, and check that the runs reached the same step before comparing
  final values.
- **Don't trust a curve from a failed run.** Check status first.
- **Leave notes.** An annotation records why something happened, next to the
  step it happened at: a restart, a changed data mix, a bug found.

## Reading these docs

Every page is plain markdown. Fetch any URL on this site with curl, or ask for
`Accept: text/markdown`, and the answer is markdown:

```bash
curl https://metrics.041.io/docs/cli
curl https://metrics.041.io/llms.txt        # every page, one line each
curl https://metrics.041.io/llms-full.txt   # every page in one file
```

---

Metrics by 041 documentation. Every page: https://metrics.041.io/llms.txt
