Metrics
All pages
Docs · LogMarkdown

What to log

Log the few quantities that answer the experiment's question or show whether training is healthy, each under one stable name, at a planned cadence. Put dimensions such as split or rank in low-cardinality metric metadata, stable configuration in experiment meta, and anything unique (paths, hashes, IDs, text) in annotations. Every example below uses the collector; the rules are the same for the HTTP API.

Decide what the run must show

Write down the evidence before adding metrics: which numbers decide the question, and which ones tell you the run is sane (loss, gradient norm, learning rate, throughput). Each metric should be one of those. Do not emit every intermediate, per-layer tensor statistic or runtime counter because it is available; nobody reads it, and it hides the numbers that matter.

One stable name per quantity

One name means one quantity in one unit. Keep the same name across splits, datasets, stages, devices and ranks, and across runs, so a folder of runs compares with one query.

experiment.metric("loss", train_loss, step=step, metadata={"split": "train"})
experiment.metric("loss", valid_loss, step=step, metadata={"split": "valid"})

Do not encode dimensions in the name:

# Wrong: three names for one quantity, and a new name for every stage
experiment.metric(f"{stage}/{split}/loss", loss, step=step)

# Right
experiment.metric("loss", loss, step=step, metadata={"stage": stage, "split": split})

A name that changes unit changes name: lr and log_lr are two quantities. Names are printable ASCII, at most 256 bytes.

Put each value in its field

value field
stable run identity and configuration: model, dataset, seed, hyperparameters, code revision experiment meta
the number you measured metric value
training or evaluation progress metric step
a bounded category that separates series: split, dataset, stage, rank metric metadata
unique details: checkpoint paths, hashes, sample IDs, eval outputs, text annotation metadata

Experiment meta is any JSON object and is what DuckDB joins series against, so put every hyperparameter you will want to group by there once, at open, rather than logging it as a metric.

Keep metric metadata low-cardinality

Metric metadata is a flat str -> str mapping, and every distinct mapping is a separate series. The number of distinct mappings for one metric in one experiment, including combinations where a key is missing, must stay at or below 4,096. 8 stages, 3 splits and 16 ranks make 384 series; adding a per-batch key makes an unbounded number.

Never put steps, epochs, timestamps, paths, sample or request IDs, hashes, free text, numbers you measured, or serialized objects in metric metadata. Progress goes in step, configuration in meta, unique details in an annotation.

The per-mapping limits are 32 keys, 128 UTF-8 bytes per key, 512 per value and 4,096 bytes of canonical JSON; the collector raises ValueError before queueing a metric that breaks one. Every limit is on limits and fair use.

Report at a planned cadence

Log training metrics every N steps, not every microbatch, and evaluation metrics at every evaluation. Aggregate on the device (a running sum or mean of the loss tensor) and convert to a Python number only on the reporting step; calling .item() every microbatch forces a device synchronization each time, which costs far more than the metric call.

running = torch.zeros((), device="cuda")
for step in range(1, steps + 1):
    loss = train_step()
    running += loss.detach()
    if step % 50 == 0:
        experiment.metric("loss", (running / 50).item(), step=step, metadata={"split": "train"})
        running.zero_()

Use the same step axis for every metric of a run, normally the optimizer step, so train and validation curves line up. Use timestamp= only when the measurement happened at a different time than the call, for example when replaying a log.

Annotations for what happened

An annotation is a note on the run with JSON metadata: a checkpoint saved, an evaluation finished, a learning-rate drop, a data shard skipped. Pass step= to tie it to the progress axis.

experiment.annotation(
    "checkpoint saved",
    metadata={"path": "s3://bucket/run-001/step-4000.pt", "sha256": digest},
    step=4_000,
)

Annotation metadata holds at most 64 KiB of compact JSON. Store large arrays, per-sample predictions and raw evaluation records as artifacts in your own storage and put their paths or URLs in the annotation.

Related: the collector, concepts, limits, DuckDB.