# Python API client

`syvain-metrics-api-client` is a small synchronous Python client for the v2 API.
Use it from scripts and agents to find and organize folders and experiments, to
read one experiment's record, catalog, annotations and series, and to send any
organization event the API accepts. For analysis across many runs use
[DuckDB](https://metrics.041.io/docs/duckdb.md); to record a training job use the [collector](https://metrics.041.io/docs/collector.md).

## Which tool

| task                                                           | use                                                                |
| -------------------------------------------------------------- | ------------------------------------------------------------------ |
| log metrics from a training or evaluation job                  | the [collector](https://metrics.041.io/docs/collector.md): background queue, batching, retries |
| tables, joins, plots over a folder of runs                     | [DuckDB](https://metrics.041.io/docs/duckdb.md): SQL with an on-disk cache                     |
| organize folders, move experiments, read one run from a script | this client                                                        |
| the same from a shell, or a quick chart                        | the [command line](https://metrics.041.io/docs/cli.md)                                         |
| backfills, tests, events no tool covers                        | this client's write session and `send()`                           |

## Install and connect

```bash
uv add syvain-metrics-api-client
```

```python
from syvain_metrics_api_client import Metrics

with Metrics() as metrics:
    print(metrics.auth_status().organization)
```

The package needs CPython 3.14.
`Metrics(api_key=None, *, host=None, timeout=60.0)` resolves credentials like
the [command line](https://metrics.041.io/docs/cli.md). The API key comes from `api_key`, then
`SYVAIN_METRICS_API_KEY`, then the login saved by `syvain-metrics auth login`.
The host comes from `host`, then `SYVAIN_METRICS_HOST`, then that login, then
`https://metrics.syvain.com`; `https://metrics.041.io` is the same service. With
no key anywhere the constructor raises `MetricsAuthError`. `timeout` is seconds
per request. `Metrics` is a context manager that closes its HTTP session;
`close()` does the same.

`auth_status()` checks the key and returns its `kind` and `organization`.

## Folders and experiments

A folder is its absolute path or its id; `/` is the root. An experiment is its
id or its slug.

```python
mamba = metrics.make_folder("/models/mamba")       # creates missing parents
mamba = metrics.folder("/models/mamba")            # MetricsNotFound when missing
mamba.folders()                                    # direct children
mamba.experiments()                                # direct members
mamba.experiments("run-*", recursive=True)         # slug glob over the subtree
metrics.experiments("lr-sweep-*")                  # slug glob over the organization

exp = metrics.experiment("mamba-run-001")
exp = exp.move("/archive")
mamba = mamba.rename("mamba-v2")
mamba = mamba.move("/")
```

`Folder` has `id` (`None` for the root), `name`, `path`, `parent_id` and
`parent`. `Experiment` has `id`, `slug`, `folder_id`, `folder_path` and
`folder`. Both are immutable values: `rename()` and `move()` return a new
handle.

Lookups use a snapshot of the folder tree, read on first use. Changes made
through the client drop it. A `folder()` or `experiment()` lookup that misses
reads the tree again once before raising `MetricsNotFound`. Listings such as
`folders()` and `experiments()` use the snapshot as it is, so call
`metrics.refresh()` to see what was created elsewhere.

## Reading one experiment

```python
info = exp.info()          # ExperimentInfo
catalog = exp.catalog()    # list[CatalogSeries]
notes = exp.annotations()  # list[Annotation], newest first

for series in exp.series("loss", where={"split": "valid"}):
    print(series.metadata, series.steps[-1], series.values[-1])
```

| method                                               | returns                                                                                                                                                               |
| ---------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `info()`                                             | `id`, `slug`, `description`, `meta`, `status`, `error` (`message`, `meta`), and `created_at`, `updated_at`, `started_at`, `done_at`, `last_event_at` as UTC datetimes |
| `catalog()`                                          | one `CatalogSeries` per series: `name` and `metadata`, each key with the tuple of values seen                                                                         |
| `annotations()`                                      | `Annotation`s with `id`, `text`, `meta`, `created_at`; at most 100,000                                                                                                |
| `series(name=None, *, where=None, has=(), x="step")` | one `Series` per metadata partition                                                                                                                                   |

`series()` takes one name, several, or `None` for every series of the
experiment. `where` keeps partitions whose metadata equals each given value,
`has` keeps partitions that carry each given key, and `x="timestamp"` orders
points by time instead of step. Results follow the order of the names you pass,
or name order for `None`. Each streaming request reads up to 1,000 series; more
names make more requests.

A `Series` has `experiment_id`, `name`, `metadata` and three aligned tuples,
`steps`, `timestamps_ms` and `values`: the values at one index are one point,
sorted ascending by the `x` axis. A step is `None` for points stored without
one. `series()` returns only after the stream completed and its counts matched,
so it never returns partial data.

## An analysis script

Final and best validation loss for every run of a sweep, with its learning rate
from `meta`, in plain Python:

```python
from syvain_metrics_api_client import Metrics

with Metrics() as metrics:
    rows = []
    for exp in metrics.experiments("lr-sweep-*"):
        info = exp.info()
        if info.status != "done":
            continue
        partitions = exp.series("loss", where={"split": "valid"})
        if len(partitions) != 1 or not partitions[0].values:
            continue  # no validation loss, or split by another key too
        valid = partitions[0]
        best = min(valid.values)
        rows.append((
            exp.slug,
            info.meta.get("lr"),
            valid.steps[-1],
            valid.values[-1],
            best,
            valid.steps[valid.values.index(best)],
        ))

rows.sort(key=lambda row: row[3])
print(f"{'run':<24} {'lr':>9} {'step':>7} {'final':>8} {'best':>8} {'at':>7}")
for slug, lr, step, final, best, best_step in rows:
    print(f"{slug:<24} {lr!s:>9} {step!s:>7} {final:8.4f} {best:8.4f} {best_step!s:>7}")
```

The script sends two requests per run, one record read and one series stream,
plus one folder tree read for the whole script. Past a few dozen runs the same
table is one [DuckDB](https://metrics.041.io/docs/duckdb.md#example-queries) query, and cached.

## Write session

```python
with metrics.open_experiment(
    "backfill-001", description="Imported run", meta={"seed": 7}, folder="/imports"
) as run:
    run.start()
    run.metric("loss", 0.25, step=1, metadata={"split": "train"})
    run.annotate("imported from wandb", meta={"source": "run/abc123"})
    run.done()
```

`open_experiment(slug, *, description="", meta=None, folder=None)` creates or
reopens the experiment and returns an `ExperimentRun` holding its experiment
token. Reopening an existing slug replaces its description and meta. An
experiment at the root is moved into `folder`; one already in another folder
stays there.

| method                                                           | request                                       |
| ---------------------------------------------------------------- | --------------------------------------------- |
| `start()`, `done()`                                              | lifecycle; not retried                        |
| `error(message, *, meta=None)`                                   | marks the run failed; not retried             |
| `annotate(text, *, meta=None)`                                   | returns the annotation id; not retried        |
| `metric(name, value, *, step, metadata=None, timestamp_ms=None)` | one point                                     |
| `metrics(points)`                                                | `MetricPoint`s, up to 1,000 per request       |
| `close()`                                                        | revokes the token; called by the `with` block |

Each call is one acknowledged request. The session refreshes its token five
minutes before it expires. A token left idle past its 24 hour expiry cannot be
refreshed; open the experiment again. A non-finite metric value raises
`ValueError` before anything is sent. For real training jobs use the
[collector](https://metrics.041.io/docs/collector.md), which never blocks the loop.

## send()

```python
reply = metrics.send("syvain.v2.query.experiments", {"experimentIds": [exp.id]})
```

`send(event, data)` sends one organization event to `/api/v2/events` and returns
the reply's `data`. It never retries and drops the tree snapshot, since the
event may change the tree. Every event and its data are listed in
[API events](https://metrics.041.io/docs/api-events.md).

## Errors and retries

Every failure is a `MetricsError` with `code`, `message`, `status`, `retryable`
and `details`. `code` is the API's error code, such as `folder_already_exists`,
or one of the client's `http_error`, `transport_error` and `invalid_response`.
`MetricsNotFound`, also a `LookupError`, reports a reference that did not
resolve. `MetricsAuthError` reports missing credentials or an invalid host.

Reads (the folder tree, `info()`, `catalog()`, `annotations()`, `series()`,
`auth_status()`) and metric writes retry transport failures, HTTP 429 and 5xx up
to four attempts, waiting 0.5, 1 and 2 seconds. A retried metric keeps its
original message id, which the API deduplicates, and only the failed points are
resent. Every other write fails on the first error, because repeating it can
have side effects: folder creation, moves, renames, `open_experiment`, lifecycle
events (a replayed `done` moves `done_at`), annotations and `send()`.

Related: [DuckDB](https://metrics.041.io/docs/duckdb.md), [the collector](https://metrics.041.io/docs/collector.md), [HTTP API](https://metrics.041.io/docs/api.md),
[command line](https://metrics.041.io/docs/cli.md).

---

Metrics by 041 documentation. Every page: https://metrics.041.io/llms.txt
