All pages
What to log
Log the few quantities that answer the experiment's question or show whether
training is healthy, each under one stable name, at a planned cadence. Put
dimensions such as split or rank in low-cardinality metric metadata, stable
configuration in experiment meta, and anything unique (paths, hashes, IDs,
text) in annotations. Every example below uses the collector;
the rules are the same for the HTTP API.
Decide what the run must show
Write down the evidence before adding metrics: which numbers decide the question, and which ones tell you the run is sane (loss, gradient norm, learning rate, throughput). Each metric should be one of those. Do not emit every intermediate, per-layer tensor statistic or runtime counter because it is available; nobody reads it, and it hides the numbers that matter.
One stable name per quantity
One name means one quantity in one unit. Keep the same name across splits, datasets, stages, devices and ranks, and across runs, so a folder of runs compares with one query.
experiment.metric("loss", train_loss, step=step, metadata={"split": "train"})
experiment.metric("loss", valid_loss, step=step, metadata={"split": "valid"})Do not encode dimensions in the name:
# Wrong: three names for one quantity, and a new name for every stage
experiment.metric(f"{stage}/{split}/loss", loss, step=step)
# Right
experiment.metric("loss", loss, step=step, metadata={"stage": stage, "split": split})A name that changes unit changes name: lr and log_lr are two quantities.
Names are printable ASCII, at most 256 bytes.
Put each value in its field
| value | field |
|---|---|
| stable run identity and configuration: model, dataset, seed, hyperparameters, code revision | experiment meta |
| the number you measured | metric value |
| training or evaluation progress | metric step |
| a bounded category that separates series: split, dataset, stage, rank | metric metadata |
| unique details: checkpoint paths, hashes, sample IDs, eval outputs, text | annotation metadata |
Experiment meta is any JSON object and is what DuckDB joins
series against, so put every hyperparameter you will want to group by there
once, at open, rather than logging it as a metric.
Keep metric metadata low-cardinality
Metric metadata is a flat str -> str mapping, and every distinct mapping is a
separate series. The number of distinct mappings for one metric in one
experiment, including combinations where a key is missing, must stay at or below
4,096. 8 stages, 3 splits and 16 ranks make 384 series; adding a per-batch key
makes an unbounded number.
Never put steps, epochs, timestamps, paths, sample or request IDs, hashes, free
text, numbers you measured, or serialized objects in metric metadata. Progress
goes in step, configuration in meta, unique details in an annotation.
The per-mapping limits are 32 keys, 128 UTF-8 bytes per key, 512 per value and
4,096 bytes of canonical JSON; the collector raises ValueError before queueing
a metric that breaks one. Every limit is on limits and fair use.
Report at a planned cadence
Log training metrics every N steps, not every microbatch, and evaluation metrics
at every evaluation. Aggregate on the device (a running sum or mean of the loss
tensor) and convert to a Python number only on the reporting step; calling
.item() every microbatch forces a device synchronization each time, which
costs far more than the metric call.
running = torch.zeros((), device="cuda")
for step in range(1, steps + 1):
loss = train_step()
running += loss.detach()
if step % 50 == 0:
experiment.metric("loss", (running / 50).item(), step=step, metadata={"split": "train"})
running.zero_()Use the same step axis for every metric of a run, normally the optimizer step,
so train and validation curves line up. Use timestamp= only when the
measurement happened at a different time than the call, for example when
replaying a log.
Annotations for what happened
An annotation is a note on the run with JSON metadata: a checkpoint saved, an
evaluation finished, a learning-rate drop, a data shard skipped. Pass step= to
tie it to the progress axis.
experiment.annotation(
"checkpoint saved",
metadata={"path": "s3://bucket/run-001/step-4000.pt", "sha256": digest},
step=4_000,
)Annotation metadata holds at most 64 KiB of compact JSON. Store large arrays, per-sample predictions and raw evaluation records as artifacts in your own storage and put their paths or URLs in the annotation.
Related: the collector, concepts, limits, DuckDB.