Metrics
All pages
Docs · LogMarkdown

The collector

syvain-metrics-collector is the Python package a training, evaluation or benchmark job logs with. Collector opens experiments, metric() and annotation() queue values without waiting on the network, a Rust core batches and delivers them in the background, and one flush_or_raise() at the end fails the job if anything was lost. To read data back, use the command line, DuckDB or the Python API client.

Install

uv add syvain-metrics-collector

The package needs CPython 3.14. It ships compiled abi3 wheels for Linux x86_64 and aarch64 (manylinux 2.28) and macOS arm64. The queue, batching, retries and HTTP delivery live in the Rust core, so there is no pure Python fallback: on any other platform the install fails.

A complete job

from syvain_metrics_collector import Collector

collector = Collector()  # reads SYVAIN_METRICS_API_KEY
experiment = collector.experiment(
    slug="mamba-run-001",
    description="Baseline mamba, 1k steps",
    folder_id="019e3bbe-dfe2-7c15-9b2e-ba2cf0f6d318",
    meta={"model": "mamba", "seed": 7, "config": {"lr": 3e-4, "batch_size": 32}},
)

with experiment.run():
    for step in range(1, 1_001):
        loss = train_step()
        if step % 10 == 0:
            experiment.metric("loss", loss, step=step, metadata={"split": "train"})
        if step % 200 == 0:
            experiment.metric("loss", evaluate(), step=step, metadata={"split": "valid"})
    experiment.annotation(
        "checkpoint saved",
        metadata={"path": "s3://bucket/mamba-run-001/step-1000.pt"},
        step=1_000,
    )

experiment.flush_or_raise()

What to name things and which field a value belongs in is on what to log.

Collector

Collector(api_key=None, *, ...) checks the API key against the API, and the default folder when one is given, before it returns. A key the API refuses, a folder that does not exist or a host that stays unreachable after three attempts raises MetricsCollectorError. Without api_key the key comes from SYVAIN_METRICS_API_KEY; when that is unset too, the constructor raises.

argument default meaning
api_key SYVAIN_METRICS_API_KEY organization API key, ak_...; see keys
host https://metrics.syvain.com API base URL; https://metrics.041.io is the same service
folder_id None folder UUID every new experiment is placed in
max_queue_items 100_000 queue capacity; at capacity the oldest metric or annotation is evicted
max_batch_items 500 events per request, 1 to 1,000
flush_delay_seconds 0.25 how long a burst is coalesced before a batch is sent
request_timeout_seconds 10.0 per-request timeout
logger logging.getLogger("syvain.metrics") logger for dropped values and exit warnings

A value outside its range, such as max_batch_items=2000 or a folder_id that is not a UUID, raises MetricsCollectorError from the constructor.

One collector serves any number of experiments; they share its queue, its worker thread and its delivery counters.

experiment()

collector.experiment(slug, description="", meta=None, *, folder_id=None)

Opens the experiment named by slug and blocks until the API accepted it, for at most 120 seconds. An open creates the experiment, or reopens an existing slug and replaces its description and meta. Afterwards experiment.id is the experiment's id and experiment.url a link to it in the app. Calling experiment() again with the same slug on the same collector returns the same Experiment without a second open.

argument limit
slug 1 to 200 characters, unique in the organization
description at most 4,000 characters
meta any JSON object; stable configuration of the run
folder_id a folder UUID; overrides the collector's folder_id for this experiment

An experiment at the root, new or reopened, is moved into the folder; one that already lives in another folder stays where it is. A folder_id that does not exist fails the open before the experiment is created. Every delivered open also adds an annotation experiment opened, so reopens are visible in the experiment's history.

An open the API rejects, or one that keeps failing transiently for 120 seconds, raises MetricsCollectorError. A rejected open also stops the collector's worker: later events are dropped and the final flush raises.

metric()

experiment.metric(name, value, step, metadata=None, *, timestamp=None)

Queues one measurement and returns. step is required: an integer from 0 to 2^53 - 2. timestamp is seconds since the epoch, or milliseconds when the number is at that scale; without it the point is stamped when it is queued.

The call validates synchronously and raises ValueError at the call site for a name that is empty, not printable ASCII or over 256 bytes, a step out of range, a metadata value that is not a string, or metadata over its limits (32 keys, 128 bytes per key, 512 bytes per value, 4,096 bytes of canonical JSON). An invalid event never reaches the queue and is never counted as dropped. A value that is NaN or infinite is logged as a warning and discarded; it is not counted as dropped either.

annotation()

experiment.annotation(text, metadata=None, *, step=None)

Queues one annotation: a note tied to the experiment, with JSON metadata. text is 1 to 16,000 characters. metadata is at most 65,536 UTF-8 bytes of compact JSON, counting the step key the API stores it under when you pass step. Oversized input raises ValueError at the call site; nothing is truncated. Put paths or URLs of large artifacts in the metadata instead of the data itself.

run()

with experiment.run():
    ...

Records the lifecycle: running on entry, done on a clean exit. When the block raises anything, KeyboardInterrupt and SystemExit included, it records error with the exception's type and message (cut to fit the API's 8,000 characters) and re-raises the exception. Recording the error never replaces the exception being raised.

flush_or_raise()

result = experiment.flush_or_raise(timeout=60.0)

Waits until every event the collector accepted before the call is delivered, then returns a FlushResult with pending, delivered, dropped, failed and error. It raises MetricsCollectorError when events are still pending at the timeout, when the worker stopped after a permanent failure, or when any event was dropped earlier in the collector's life: evicted at queue capacity, or rejected by the API in a background batch. Drop counts never reset, so a later clean drain does not hide an earlier loss. Events queued after the call started, by other threads or experiments, are not waited for.

Call it once, after the run() block. Calling it from the training loop stalls training on the network and adds nothing: delivery already runs in the background.

Exit drain

At interpreter exit the collector waits up to 30 seconds for queued events, logs a warning naming the counts if events are still pending or were dropped, then stops its worker within what is left of those 30 seconds and revokes its experiment tokens on a best-effort basis. The exit drain never raises: a job that must fail on lost data calls flush_or_raise().

Delivery

The worker sends batches of up to max_batch_items events, grouped per experiment and in queue order. Connection errors and HTTP 408, 425, 429 and 5xx replies are retried with backoff from 500 ms doubling to 10 s, until the collector closes. A metric whose own reply is a 429 or 5xx error is resent with its original message id, which the API deduplicates; a lifecycle event or annotation is not resent alone, because it would land after later events. Any other rejection drops the event and counts it.

Experiment tokens expire after 24 hours. The worker refreshes a token five minutes before it expires and reopens the slug with its original description and meta when the token is already gone, so a job that runs for days keeps delivering. A refresh the API refuses because the organization key itself is no longer valid stops the worker.

Measured with JsonlCollector on a GCP n2-standard-8 Linux VM (Xeon at 2.8 GHz, CPython 3.14), metric() costs about 11 µs per call: 1,000,000 points were queued and written to the file in 10.7 s with none dropped, the median of five trials. That measures the Python call and the Rust queue; delivery over the network depends on the API and the connection.

JsonlCollector and NoopCollector

Both have the same experiment(), metric(), annotation(), run() and flush_or_raise() as Collector and need no API key.

from syvain_metrics_collector import JsonlCollector, NoopCollector

local = JsonlCollector("runs/metrics.jsonl")
silent = NoopCollector()

JsonlCollector(path="metrics.jsonl", *, max_queue_items, max_batch_items, flush_delay_seconds, logger) appends one JSON object per event to path, creating parent directories. Records have type experiment, status, metric or annotation, the slug and timestamp_ms. Its experiments have id equal to the slug and a local:// URL.

NoopCollector(*, logger=None) does no IO. It validates every event with the same rules, so a test still sees ValueError for bad input, and reports everything as delivered.

Errors

error raised by when
ValueError experiment(), metric(), annotation() the event breaks a limit; nothing is queued
MetricsCollectorError Collector(...) missing or refused key, unknown folder, bad argument, API unreachable
MetricsCollectorError experiment() the open was rejected or timed out
MetricsCollectorError flush_or_raise() pending at timeout, dropped events, or a stopped worker

MetricsCollectorError is a RuntimeError. Its message carries the counts and the most recent delivery error.

Related: what to log, limits, concepts, HTTP API.