# Series storage

Metric points live in Syvain Series DB, a storage system we built for one job:
experiment metrics. General-purpose time-series databases are tuned for
monitoring, with rolling windows, downsampling and dashboards that refresh every
few seconds. Training runs are different: bursts of writes while a run is live,
then years of reads of a run that never changes again, sliced by metadata such
as split, layer or rank. Series DB is shaped around that, to keep every point
while staying fast and cheap to run.

## One series, one owner

Each metric of each experiment is its own series, with its own small database on
Cloudflare's edge. A write is committed there before the API answers, so an
acknowledged point is durable, and a retried request never appends a point
twice. Series never share a lock, so one busy run does not slow another.

Inside a series, every distinct metadata map is a partition: `loss` with
`{"split": "train"}` and `loss` with `{"split": "valid"}` are two partitions of
one series and two lines on a chart. See [Concepts](https://metrics.041.io/docs/concepts.md).

## Live runs and finished runs

While a run is live, its points move in the background into compact, immutable
columnar files in object storage, and ingest never waits on that. A read of a
live run sees one consistent snapshot of the files and the newest points
together, so a running experiment is queryable at any moment.

Once a run goes quiet, its series is sealed: compacted into as few files as its
data allows. From then on a read touches only object storage. Finished runs,
which is most of what anyone reads, cost almost nothing to keep and wake no
database to query.

Research metadata makes many small partitions, one per layer, head or evaluation
set. Series DB packs those together instead of writing thousands of tiny files,
and a query fetches and decodes only the partitions it asked for.

## Reading at the scale of a sweep

A whole experiment streams back as one response rather than one request per
metric, and the catalog of an experiment (its metric names with every metadata
value) is read from summaries kept next to the data, not by scanning it. That is
what lets the [DuckDB extension](https://metrics.041.io/docs/duckdb.md) pull a folder of runs in one pass,
and the app draw a sweep's charts as you scroll.

## What it means for you

- **Lossless.** Every accepted point is kept at full resolution: a float64
  value, an integer step and a millisecond timestamp. Nothing is downsampled or
  aggregated, so what you read is what you logged.
- **Retries are safe.** A metric event retried with the same `messageId` is
  appended once.
- **Read while it writes.** A running experiment is queryable at any moment,
  without waiting for anything to be flushed.
- **Finished runs are cheap.** A sealed run is immutable files, so reading an
  old run never contends with the runs being trained.
- **Caches know what changed.** Every experiment carries revision counters, and
  the [DuckDB](https://metrics.041.io/docs/duckdb.md) extension downloads again only the experiments whose
  counter moved.

The limits on keys, metadata and batches are in
[Limits and fair use](https://metrics.041.io/docs/limits.md).

---

Metrics by 041 documentation. Every page: https://metrics.041.io/llms.txt
