# The collector

`syvain-metrics-collector` is the Python package a training, evaluation or
benchmark job logs with. `Collector` opens experiments, `metric()` and
`annotation()` queue values without waiting on the network, a Rust core batches
and delivers them in the background, and one `flush_or_raise()` at the end fails
the job if anything was lost. To read data back, use the [command line](https://metrics.041.io/docs/cli.md),
[DuckDB](https://metrics.041.io/docs/duckdb.md) or the [Python API client](https://metrics.041.io/docs/python-client.md).

## Install

```bash
uv add syvain-metrics-collector
```

The package needs CPython 3.14. It ships compiled abi3 wheels for Linux x86_64
and aarch64 (manylinux 2.28) and macOS arm64. The queue, batching, retries and
HTTP delivery live in the Rust core, so there is no pure Python fallback: on any
other platform the install fails.

## A complete job

```python
from syvain_metrics_collector import Collector

collector = Collector()  # reads SYVAIN_METRICS_API_KEY
experiment = collector.experiment(
    slug="mamba-run-001",
    description="Baseline mamba, 1k steps",
    folder_id="019e3bbe-dfe2-7c15-9b2e-ba2cf0f6d318",
    meta={"model": "mamba", "seed": 7, "config": {"lr": 3e-4, "batch_size": 32}},
)

with experiment.run():
    for step in range(1, 1_001):
        loss = train_step()
        if step % 10 == 0:
            experiment.metric("loss", loss, step=step, metadata={"split": "train"})
        if step % 200 == 0:
            experiment.metric("loss", evaluate(), step=step, metadata={"split": "valid"})
    experiment.annotation(
        "checkpoint saved",
        metadata={"path": "s3://bucket/mamba-run-001/step-1000.pt"},
        step=1_000,
    )

experiment.flush_or_raise()
```

What to name things and which field a value belongs in is on
[what to log](https://metrics.041.io/docs/logging-guide.md).

## Collector

`Collector(api_key=None, *, ...)` checks the API key against the API, and the
default folder when one is given, before it returns. A key the API refuses, a
folder that does not exist or a host that stays unreachable after three attempts
raises `MetricsCollectorError`. Without `api_key` the key comes from
`SYVAIN_METRICS_API_KEY`; when that is unset too, the constructor raises.

| argument                  | default                               | meaning                                                                |
| ------------------------- | ------------------------------------- | ---------------------------------------------------------------------- |
| `api_key`                 | `SYVAIN_METRICS_API_KEY`              | organization API key, `ak_...`; see [keys](https://metrics.041.io/docs/organizations.md)           |
| `host`                    | `https://metrics.syvain.com`          | API base URL; `https://metrics.041.io` is the same service             |
| `folder_id`               | `None`                                | folder UUID every new experiment is placed in                          |
| `max_queue_items`         | `100_000`                             | queue capacity; at capacity the oldest metric or annotation is evicted |
| `max_batch_items`         | `500`                                 | events per request, 1 to 1,000                                         |
| `flush_delay_seconds`     | `0.25`                                | how long a burst is coalesced before a batch is sent                   |
| `request_timeout_seconds` | `10.0`                                | per-request timeout                                                    |
| `logger`                  | `logging.getLogger("syvain.metrics")` | logger for dropped values and exit warnings                            |

A value outside its range, such as `max_batch_items=2000` or a `folder_id` that
is not a UUID, raises `MetricsCollectorError` from the constructor.

One collector serves any number of experiments; they share its queue, its worker
thread and its delivery counters.

## experiment()

```python
collector.experiment(slug, description="", meta=None, *, folder_id=None)
```

Opens the experiment named by `slug` and blocks until the API accepted it, for
at most 120 seconds. An open creates the experiment, or reopens an existing slug
and replaces its description and meta. Afterwards `experiment.id` is the
experiment's id and `experiment.url` a link to it in the app. Calling
`experiment()` again with the same slug on the same collector returns the same
`Experiment` without a second open.

| argument      | limit                                                                    |
| ------------- | ------------------------------------------------------------------------ |
| `slug`        | 1 to 200 characters, unique in the organization                          |
| `description` | at most 4,000 characters                                                 |
| `meta`        | any JSON object; stable configuration of the run                         |
| `folder_id`   | a folder UUID; overrides the collector's `folder_id` for this experiment |

An experiment at the root, new or reopened, is moved into the folder; one that
already lives in another folder stays where it is. A `folder_id` that does not
exist fails the open before the experiment is created. Every delivered open also
adds an annotation `experiment opened`, so reopens are visible in the
experiment's history.

An open the API rejects, or one that keeps failing transiently for 120 seconds,
raises `MetricsCollectorError`. A rejected open also stops the collector's
worker: later events are dropped and the final flush raises.

## metric()

```python
experiment.metric(name, value, step, metadata=None, *, timestamp=None)
```

Queues one measurement and returns. `step` is required: an integer from 0 to
2^53 - 2. `timestamp` is seconds since the epoch, or milliseconds when the
number is at that scale; without it the point is stamped when it is queued.

The call validates synchronously and raises `ValueError` at the call site for a
name that is empty, not printable ASCII or over 256 bytes, a step out of range,
a metadata value that is not a string, or metadata over its limits (32 keys, 128
bytes per key, 512 bytes per value, 4,096 bytes of canonical JSON). An invalid
event never reaches the queue and is never counted as dropped. A value that is
NaN or infinite is logged as a warning and discarded; it is not counted as
dropped either.

## annotation()

```python
experiment.annotation(text, metadata=None, *, step=None)
```

Queues one annotation: a note tied to the experiment, with JSON metadata. `text`
is 1 to 16,000 characters. `metadata` is at most 65,536 UTF-8 bytes of compact
JSON, counting the `step` key the API stores it under when you pass `step`.
Oversized input raises `ValueError` at the call site; nothing is truncated. Put
paths or URLs of large artifacts in the metadata instead of the data itself.

## run()

```python
with experiment.run():
    ...
```

Records the lifecycle: `running` on entry, `done` on a clean exit. When the
block raises anything, `KeyboardInterrupt` and `SystemExit` included, it records
`error` with the exception's type and message (cut to fit the API's 8,000
characters) and re-raises the exception. Recording the error never replaces the
exception being raised.

## flush_or_raise()

```python
result = experiment.flush_or_raise(timeout=60.0)
```

Waits until every event the collector accepted before the call is delivered,
then returns a `FlushResult` with `pending`, `delivered`, `dropped`, `failed`
and `error`. It raises `MetricsCollectorError` when events are still pending at
the timeout, when the worker stopped after a permanent failure, or when any
event was dropped earlier in the collector's life: evicted at queue capacity, or
rejected by the API in a background batch. Drop counts never reset, so a later
clean drain does not hide an earlier loss. Events queued after the call started,
by other threads or experiments, are not waited for.

Call it once, after the `run()` block. Calling it from the training loop stalls
training on the network and adds nothing: delivery already runs in the
background.

## Exit drain

At interpreter exit the collector waits up to 30 seconds for queued events, logs
a warning naming the counts if events are still pending or were dropped, then
stops its worker within what is left of those 30 seconds and revokes its
experiment tokens on a best-effort basis. The exit drain never raises: a job
that must fail on lost data calls `flush_or_raise()`.

## Delivery

The worker sends batches of up to `max_batch_items` events, grouped per
experiment and in queue order. Connection errors and HTTP 408, 425, 429 and 5xx
replies are retried with backoff from 500 ms doubling to 10 s, until the
collector closes. A metric whose own reply is a 429 or 5xx error is resent with
its original message id, which the API deduplicates; a lifecycle event or
annotation is not resent alone, because it would land after later events. Any
other rejection drops the event and counts it.

Experiment tokens expire after 24 hours. The worker refreshes a token five
minutes before it expires and reopens the slug with its original description and
meta when the token is already gone, so a job that runs for days keeps
delivering. A refresh the API refuses because the organization key itself is no
longer valid stops the worker.

Measured with `JsonlCollector` on a GCP n2-standard-8 Linux VM (Xeon at 2.8 GHz,
CPython 3.14), `metric()` costs about 11 µs per call: 1,000,000 points were
queued and written to the file in 10.7 s with none dropped, the median of five
trials. That measures the Python call and the Rust queue; delivery over the
network depends on the API and the connection.

## JsonlCollector and NoopCollector

Both have the same `experiment()`, `metric()`, `annotation()`, `run()` and
`flush_or_raise()` as `Collector` and need no API key.

```python
from syvain_metrics_collector import JsonlCollector, NoopCollector

local = JsonlCollector("runs/metrics.jsonl")
silent = NoopCollector()
```

`JsonlCollector(path="metrics.jsonl", *, max_queue_items, max_batch_items, flush_delay_seconds, logger)`
appends one JSON object per event to `path`, creating parent directories.
Records have `type` `experiment`, `status`, `metric` or `annotation`, the `slug`
and `timestamp_ms`. Its experiments have `id` equal to the slug and a `local://`
URL.

`NoopCollector(*, logger=None)` does no IO. It validates every event with the
same rules, so a test still sees `ValueError` for bad input, and reports
everything as delivered.

## Errors

| error                   | raised by                                  | when                                                                  |
| ----------------------- | ------------------------------------------ | --------------------------------------------------------------------- |
| `ValueError`            | `experiment()`, `metric()`, `annotation()` | the event breaks a limit; nothing is queued                           |
| `MetricsCollectorError` | `Collector(...)`                           | missing or refused key, unknown folder, bad argument, API unreachable |
| `MetricsCollectorError` | `experiment()`                             | the open was rejected or timed out                                    |
| `MetricsCollectorError` | `flush_or_raise()`                         | pending at timeout, dropped events, or a stopped worker               |

`MetricsCollectorError` is a `RuntimeError`. Its message carries the counts and
the most recent delivery error.

Related: [what to log](https://metrics.041.io/docs/logging-guide.md), [limits](https://metrics.041.io/docs/limits.md),
[concepts](https://metrics.041.io/docs/concepts.md), [HTTP API](https://metrics.041.io/docs/api.md).

---

Metrics by 041 documentation. Every page: https://metrics.041.io/llms.txt
