All pages
Python API client
syvain-metrics-api-client is a small synchronous Python client for the v2 API.
Use it from scripts and agents to find and organize folders and experiments, to
read one experiment's record, catalog, annotations and series, and to send any
organization event the API accepts. For analysis across many runs use
DuckDB; to record a training job use the collector.
Which tool
| task | use |
|---|---|
| log metrics from a training or evaluation job | the collector: background queue, batching, retries |
| tables, joins, plots over a folder of runs | DuckDB: SQL with an on-disk cache |
| organize folders, move experiments, read one run from a script | this client |
| the same from a shell, or a quick chart | the command line |
| backfills, tests, events no tool covers | this client's write session and send() |
Install and connect
uv add syvain-metrics-api-clientfrom syvain_metrics_api_client import Metrics
with Metrics() as metrics:
print(metrics.auth_status().organization)The package needs CPython 3.14.
Metrics(api_key=None, *, host=None, timeout=60.0) resolves credentials like
the command line. The API key comes from api_key, then
SYVAIN_METRICS_API_KEY, then the login saved by syvain-metrics auth login.
The host comes from host, then SYVAIN_METRICS_HOST, then that login, then
https://metrics.syvain.com; https://metrics.041.io is the same service. With
no key anywhere the constructor raises MetricsAuthError. timeout is seconds
per request. Metrics is a context manager that closes its HTTP session;
close() does the same.
auth_status() checks the key and returns its kind and organization.
Folders and experiments
A folder is its absolute path or its id; / is the root. An experiment is its
id or its slug.
mamba = metrics.make_folder("/models/mamba") # creates missing parents
mamba = metrics.folder("/models/mamba") # MetricsNotFound when missing
mamba.folders() # direct children
mamba.experiments() # direct members
mamba.experiments("run-*", recursive=True) # slug glob over the subtree
metrics.experiments("lr-sweep-*") # slug glob over the organization
exp = metrics.experiment("mamba-run-001")
exp = exp.move("/archive")
mamba = mamba.rename("mamba-v2")
mamba = mamba.move("/")Folder has id (None for the root), name, path, parent_id and
parent. Experiment has id, slug, folder_id, folder_path and
folder. Both are immutable values: rename() and move() return a new
handle.
Lookups use a snapshot of the folder tree, read on first use. Changes made
through the client drop it. A folder() or experiment() lookup that misses
reads the tree again once before raising MetricsNotFound. Listings such as
folders() and experiments() use the snapshot as it is, so call
metrics.refresh() to see what was created elsewhere.
Reading one experiment
info = exp.info() # ExperimentInfo
catalog = exp.catalog() # list[CatalogSeries]
notes = exp.annotations() # list[Annotation], newest first
for series in exp.series("loss", where={"split": "valid"}):
print(series.metadata, series.steps[-1], series.values[-1])| method | returns |
|---|---|
info() |
id, slug, description, meta, status, error (message, meta), and created_at, updated_at, started_at, done_at, last_event_at as UTC datetimes |
catalog() |
one CatalogSeries per series: name and metadata, each key with the tuple of values seen |
annotations() |
Annotations with id, text, meta, created_at; at most 100,000 |
series(name=None, *, where=None, has=(), x="step") |
one Series per metadata partition |
series() takes one name, several, or None for every series of the
experiment. where keeps partitions whose metadata equals each given value,
has keeps partitions that carry each given key, and x="timestamp" orders
points by time instead of step. Results follow the order of the names you pass,
or name order for None. Each streaming request reads up to 1,000 series; more
names make more requests.
A Series has experiment_id, name, metadata and three aligned tuples,
steps, timestamps_ms and values: the values at one index are one point,
sorted ascending by the x axis. A step is None for points stored without
one. series() returns only after the stream completed and its counts matched,
so it never returns partial data.
An analysis script
Final and best validation loss for every run of a sweep, with its learning rate
from meta, in plain Python:
from syvain_metrics_api_client import Metrics
with Metrics() as metrics:
rows = []
for exp in metrics.experiments("lr-sweep-*"):
info = exp.info()
if info.status != "done":
continue
partitions = exp.series("loss", where={"split": "valid"})
if len(partitions) != 1 or not partitions[0].values:
continue # no validation loss, or split by another key too
valid = partitions[0]
best = min(valid.values)
rows.append((
exp.slug,
info.meta.get("lr"),
valid.steps[-1],
valid.values[-1],
best,
valid.steps[valid.values.index(best)],
))
rows.sort(key=lambda row: row[3])
print(f"{'run':<24} {'lr':>9} {'step':>7} {'final':>8} {'best':>8} {'at':>7}")
for slug, lr, step, final, best, best_step in rows:
print(f"{slug:<24} {lr!s:>9} {step!s:>7} {final:8.4f} {best:8.4f} {best_step!s:>7}")The script sends two requests per run, one record read and one series stream, plus one folder tree read for the whole script. Past a few dozen runs the same table is one DuckDB query, and cached.
Write session
with metrics.open_experiment(
"backfill-001", description="Imported run", meta={"seed": 7}, folder="/imports"
) as run:
run.start()
run.metric("loss", 0.25, step=1, metadata={"split": "train"})
run.annotate("imported from wandb", meta={"source": "run/abc123"})
run.done()open_experiment(slug, *, description="", meta=None, folder=None) creates or
reopens the experiment and returns an ExperimentRun holding its experiment
token. Reopening an existing slug replaces its description and meta. An
experiment at the root is moved into folder; one already in another folder
stays there.
| method | request |
|---|---|
start(), done() |
lifecycle; not retried |
error(message, *, meta=None) |
marks the run failed; not retried |
annotate(text, *, meta=None) |
returns the annotation id; not retried |
metric(name, value, *, step, metadata=None, timestamp_ms=None) |
one point |
metrics(points) |
MetricPoints, up to 1,000 per request |
close() |
revokes the token; called by the with block |
Each call is one acknowledged request. The session refreshes its token five
minutes before it expires. A token left idle past its 24 hour expiry cannot be
refreshed; open the experiment again. A non-finite metric value raises
ValueError before anything is sent. For real training jobs use the
collector, which never blocks the loop.
send()
reply = metrics.send("syvain.v2.query.experiments", {"experimentIds": [exp.id]})send(event, data) sends one organization event to /api/v2/events and returns
the reply's data. It never retries and drops the tree snapshot, since the
event may change the tree. Every event and its data are listed in
API events.
Errors and retries
Every failure is a MetricsError with code, message, status, retryable
and details. code is the API's error code, such as folder_already_exists,
or one of the client's http_error, transport_error and invalid_response.
MetricsNotFound, also a LookupError, reports a reference that did not
resolve. MetricsAuthError reports missing credentials or an invalid host.
Reads (the folder tree, info(), catalog(), annotations(), series(),
auth_status()) and metric writes retry transport failures, HTTP 429 and 5xx up
to four attempts, waiting 0.5, 1 and 2 seconds. A retried metric keeps its
original message id, which the API deduplicates, and only the failed points are
resent. Every other write fails on the first error, because repeating it can
have side effects: folder creation, moves, renames, open_experiment, lifecycle
events (a replayed done moves done_at), annotations and send().
Related: DuckDB, the collector, HTTP API, command line.