---
order: 0
subtitle: How Benchwright works. Pipeline phases, sandboxes, events, fingerprints, scorers.
---

# Benchwright Concepts

A conceptual reference for the pieces you'll see in logs, events, and
responses. Read this first if the other docs use a term and you're
not sure what it means.

For HTTP endpoints see [api.md](./api.md). For task-oriented
walkthroughs see [developer.md](./developer.md).

## Benchmarking AI models, briefly

A **benchmark** is a fixed set of problems with known answers. You
hand each problem to a model, compare the model's answer to the
reference answer, and aggregate the per-problem outcomes into a
single score. The score is the model's *measured capability* on
whatever the benchmark is designed to test: math word problems
(GSM8K), undergraduate exam questions (MMLU), code completion
(HumanEval), tool calling (BFCL), and so on.

**Why it matters.** Vendor numbers are marketing. A self-run
benchmark answers four questions vendor decks can't:

- *Is the new model actually better at the thing I care about?*
  Headline numbers (MMLU, GPQA) average across many domains; your
  workload is a slice of one of them.
- *Did a model regress between releases?* Providers retrain and
  re-quantize silently. A pinned benchmark is your regression test.
- *Which model is cheapest at acceptable quality?* Quality and cost
  trade off; without a number for both, you can't price the trade.
- *Does fine-tuning / a prompt change help or hurt?* You need an
  honest yardstick before you ship.

**The three pieces you need.**

- A **model** (the system under test, or SUT) is anything that
  takes input and produces output: an LLM API, a local model, a
  retrieval-augmented chain, an agent. Benchwright doesn't care
  what's inside it, only that an HTTP endpoint will respond.
- A **dataset** holds the problems and reference answers.
- A **scorer** runs the model against each problem and decides
  whether the response counts as correct. This is the part that's
  hard to share between projects, because real-world correctness is
  rarely exact-match: math answers need a normalizer, tool calls
  need a JSON comparison, code needs to actually execute against
  tests.

**Where Benchwright fits.** Most teams stall at the scorer: writing
one is fiddly, and once you have it, running 1k tasks in parallel,
keeping API keys out of untrusted code, capping spend, and replaying
fairly across models is its own project. Benchwright is that project,
delivered as one pipeline. Point it at a dataset (or a natural-language
description of one), pick a model, and you get a comparable score
with the full event log behind it. The rest of this doc explains the
pieces of that pipeline, starting with what one execution of it (a
**run**) is and how it's structured.

## 1. What a run is

A **run** is one execution of the Benchwright pipeline. Every run has
a stable ID (`<uuid>`) and an append-only event log. Anything you
can observe about a run (status, cost, pass rate, phase progression)
is a *projection* of that event log. The log is authoritative; the
projections are convenience.

A run's stored `status` is one of three values:

- `running`: the pipeline is still executing.
- `ok`: Materialize produced results and the run reached its terminal
  receipt.
- `error`: a phase bailed out. The partial event log is still valid.

Halting is a separate flag, not a status. A run torn down mid-flight
(balance exhausted, manual stop, CLI halt) keeps whatever `status` it
had and carries `halted = true` alongside a `halt_reason`. See §9.

API responses return all three fields plus a derived `display_status`
that folds them together: `ok`, `error`, `halted`, `killed`, `waiting`,
`paused`. Render `display_status`; it is the only one of the four that
reads correctly for a deliberate pause.

The run is the primary unit everything else hangs off: events,
artifacts, per-task results, debits, forks.

---

## 2. Pipeline phases

A natural-language run executes these phases in order. Each one
either produces a declared artifact, mutates the work sandbox (see
§4), or both.

| Phase           | What it does                                                                                                                                             | Produces                                      |
|-----------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------------------------------|
| **Find**        | Identifies the benchmark and model from the natural-language request. Locates sources (GitHub repos, HuggingFace datasets, papers, web pages).             | A source list and the resolved model identity. |
| **Fetch**       | Downloads the non-paper sources into the work sandbox, retrying alternatives on failure. Snapshots the sandbox so later phases can rewind to this state. | Fetched files on disk; a post-fetch snapshot. |
| **Analyze**     | Explores the fetched files to recover dataset metadata: structure, available subsets, sample counts, native scoring method.                              | A structured analysis artifact.              |
| **Setup**       | Installs lightweight deps (pandas, pyarrow, etc.) in the work sandbox. Catalogs reference scripts. Model infra (GPU, torch) stays out; it lives in proxy. | A post-setup snapshot.                       |
| **Attach**      | Boots the proxy sandbox, writes a `litellm.yaml` targeting the model, runs a single-request smoke test to prove the model is reachable, publishes the proxy URL. | A running proxy sandbox and its URL.         |
| **Atomize**     | Writes `task_impl.py`, a Python class with `extract`, `solve`, and `score` methods. Proves it works on sample indices before publishing. Snapshots the sandbox at exit. | `task_impl.py` (the impl, see §7); a post-atomize snapshot. |
| **Materialize** | Runs the full dataset through `task_impl.py`, streaming per-task results. Parallelized through the cluster sandbox queue up to your tier's concurrent-sandbox ceiling, with a self-tuning dispatch rate. | Per-task rows in the trace store.            |
| **Report**      | Off by default. When enabled, renders an LLM-written summary from Materialize's per-task rows. Otherwise the run page ends in a deterministic Receipt (params, compute, cost). | The report artifact, when enabled.           |

**Why Attach sits between Setup and Atomize.** Fetch, Analyze, and
Setup are model-agnostic: they pull bytes, inspect schema, and install
dependencies. Running them before Attach keeps their snapshots free of
per-user credentials, so the cache cells they populate are shareable
across every user running the same benchmark. Attach binds this user's
proxy sandbox and credentials, and runs immediately before Atomize so
`task_impl.py` is written against a model that is actually live.

`Fork` is a pre-pipeline phase that runs only on forked launches
(§10). It reads the operator's fork instructions and decides where
to rewind to and which state to override.

**Invocation numbers.** A phase can be retried inside a single run
without starting a new run. Each retry increments a 1-based
**invocation** counter for that phase. `Atomize v2` means the
scorer was rewritten and re-submitted; `Materialize v3` means the
task sweep was run three times (usually because Atomize produced a
new impl between sweeps). The invocation number is carried by the
atomize artifacts and the Materialize result events, so the dashboard
can group results by the sweep that produced them. Phase events
themselves do not carry it; a second `PhaseStarted(phase=atomize)` on
the same run is how you spot a retry.

**Skip logic.** Two independent mechanisms decide whether a phase
actually does work.

- **Hard skips by run kind.** A harness-bridge run (`inspect`,
  `lmeval`, `lighteval`) or a product-sandbox run skips Fetch,
  Analyze, and Atomize outright, because the harness supplies its own
  dataset and its own scorer. There is nothing to download, no schema
  to discover, and no `task_impl.py` to write. Find skips whenever the
  benchmark name is already resolved, which is the case for a bridge
  task (the task *is* the identity).
- **Cross-run cache hits.** Any phase can hit the cache: if a prior
  public run recorded a matching cell in `phase_fingerprints`, the
  phase copies that run's artifact and snapshot into state and skips
  its Driver work entirely. This is what makes a repeat run of a known
  benchmark cheap, and it is keyed off the phase's input hash, not off
  anything in your launch body.

So the happy path for a re-run of a known benchmark is
`Fetch -> Setup -> Attach -> Materialize`, with Fetch and Setup
usually cache-hitting rather than re-running. Setup is never skipped
by a pin; it either cache-hits or runs.

---

## 3. Launch modes

Four bodies, one pipeline. The launcher picks which phases to skip
based on which fields you provide.

| Mode                | Required fields                       | First-run cost | Skips                  |
|---------------------|---------------------------------------|----------------|------------------------|
| Natural-language    | `request`                             | full           | nothing                |
| Pinned              | `benchmark`, `subset`, `impl`         | low            | Find, Analyze, Atomize |
| Replay              | `replay.from_run`                     | low            | same as pinned         |
| Fork                | `description` (on `POST /v1/runs/{id}/fork`) | low     | everything above the rewind point (§10) |

Rule of thumb: the first run against any new benchmark is
natural-language (it does the work of finding, analyzing, and
atomizing). Everything after is a replay.

> Pinned mode resolves `benchmark` and `subset` today: the launcher
> looks each up in the registry (by id, then by exact name) and seeds
> the launch from it. `impl` is accepted and recorded on the run as
> `metadata.pinned_impl_id`, but it is not resolved, and there is no
> server-side handling of `impl="latest:official"` even though the CLI
> advertises the flag. Pinning an impl today records your intent; it
> does not select the scorer.

---

## 4. The sandbox model

Every run spawns sandboxes on the configured compute backend, each
with a distinct role. The backend is either sandbox0 (Kubernetes pods,
the default) or E2B (managed microVMs), and it is recorded on the
trace as `backend`. The `run.cost.sandbox` rate depends on the
backend. The role is written into sandbox
`metadata.benchwright_role` and is how events, cost reconstruction,
and the dashboard tell them apart.

Backend selection has three levels, narrowest first:

1. The per-run `compute` launch field, which accepts `e2b`,
   `sandbox0`, or `default` / `auto` (omit it for your default). A
   typo is rejected at launch rather than silently falling back.
2. Your account's saved compute provider, if you configured a
   self-hosted sandbox0 endpoint under Settings → Compute.
3. The system default, which is the shared Benchwright Cloud
   sandbox0 cluster. E2B is the historical fallback and applies only
   when neither of the above is configured.

The `harness` launch field selects how the benchmark evaluation is
executed inside those sandboxes: `benchwright` (our own harness — an
Driver-authored scorer, the default), `inspect`, `lmeval`, `lighteval`,
or `harbor`.

| Role                 | Contains                                | Spawned by    | Lives for                    |
|----------------------|------------------------------------------|---------------|------------------------------|
| **`driver_sandbox`** | `port.jar` (the pipeline driver)        | RunLauncher   | the full run                 |
| **`work_sandbox`**   | Fetched sources, `task_impl.py`, tooling| AgentPipeline | the full run                 |
| **`proxy_sandbox`**  | LiteLLM, the real model credentials      | Attach        | the full run                 |
| **`task_sandbox`**   | One task's execution environment        | Materialize   | one task (fanout mode only)  |

Three more roles show up on specific run kinds: `template_probe` (a
product-sandbox run probing its base image), and `harbor_orchestrator`
plus `harbor_task_env` on harbor bridge runs.

The **driver sandbox** is where the pipeline itself executes. It
never sees model credentials.

The **work sandbox** is shared across all phases. Each phase that
changes its state re-snapshots it at exit, so later phases (and
forks) can restore to a specific mid-pipeline state without
re-running everything.

The **proxy sandbox** exists specifically to isolate model credentials.
The work sandbox receives only a fake key; the proxy forwards the
call to the real provider with the real key. You don't need to
trust a scorer's network code, because the scorer can't reach the
provider directly.

The **task sandbox** is optional. Materialize has two execution
modes:

- **Sequential**: tasks run inside the shared work sandbox. Cheaper
  and faster for lightweight scorers.
- **Fanout**: each task gets its own ephemeral task sandbox, created
  for that task and killed when it finishes. Used when the scorer
  needs isolation between tasks (state leaks, heavy side-effects)
  or parallelism beyond what the work sandbox can safely share.

The scorer does not choose. The pipeline picks the mode from the
run's infrastructure alone, in this order: a product-sandbox run with
a snapshot to restore gets fanout; a harness-bridge run gets
sequential; an E2B run gets fanout (E2B spawns fast enough that a
sandbox per task is worth it); everything else gets sequential. On
the default sandbox0 backend that means Materialize runs sequential.

**Work sandbox snapshots.** Fetch, Setup, and Atomize each snapshot
the work sandbox at exit. A snapshot is a restorable image of the
sandbox's filesystem and process state, identified by a
`snapshot_id`. Later phases (and forks) can restore to a specific
snapshot instead of re-running every phase above it. This is why
rewinding to Atomize in a fork doesn't re-download sources or
reinstall dependencies: the post-setup snapshot is the starting
point. The post-atomize snapshot is the one fanout Materialize
restores into each task sandbox, which is how every task starts from
a filesystem that already has `task_impl.py` on it.

Sandbox lifecycle is visible in the event stream as `SandboxCreated`
and `SandboxKilled` events. Filter by `metadata.benchwright_role` to
reconstruct what happened to any one role.

---

## 5. Events and projections

The event log is the source of truth. A run's event stream is a
monotonically-numbered sequence (`seq` starts at 0) of typed
records, each with a `type` and a `data` payload. The run-level
`RunLaunched` event always lands at `seq=0`.

Common types you'll see:

- `RunLaunched`: emitted by the website at `seq=0`. `PipelineStarted`
  and `PipelineFinished` are the driver's own lifecycle markers,
  emitted from inside the run.
- `PhaseStarted`, `PhaseCompleted`, `PhaseFailed`: one pair per
  phase *invocation* (see §2). `PhaseStarted` carries only `phase`;
  `PhaseCompleted` adds `duration_ms`; `PhaseFailed` adds
  `error_type`, `detail`, and `stack_trace`. None of them carry the
  invocation number, so two `PhaseStarted(phase=atomize)` records on
  the same run are how you spot a retry.
- `SandboxCreated`, `SandboxKilled`: one pair per sandbox. The role
  is in `metadata.benchwright_role`.
- `SnapshotCreated`: a work-sandbox snapshot was taken. Carries the
  `snapshot_id` a fork would target.
- `AgentTextEvent`, `ToolCallEvent`, `ToolResultEvent`: fine-grained
  driver traffic inside a phase.
- `MaterializeTaskResult`: one row per *trial* during Materialize.
  A task with `trials.plan.k = 1` produces one of these per task; a
  task with `k > 1` produces `k` of them, all stamped with the same
  `(invocation, index)` and different `trial` values (§7).
- `MaterializeTaskAggregated`: one per task once all of its trials
  have settled. Carries the row-level score the pass-rate
  projection is built from, plus the per-trial scores that produced
  it. With `k = 1` this echoes the single trial's score.

There is no cost event. Cost is accumulated straight onto the run row
(`traces.cost_*`) by a database trigger that fires as events land, so
you read a running total off the run rather than summing deltas out of
the stream.

Projections you see in API responses (`run.cost.incurred`, `run.status`,
`tasks.summary.pass_rate`) are recomputed from the log. The log is
never rewritten; projections can lag but will always converge.

Two practical consequences:

- **SSE and polling give you the same data.** Streaming is a UX
  optimization; nothing new is recorded on a stream that isn't in
  the log.
- **Failed runs are still useful.** An `error` run has the full
  log up to the failure, which is how the run page still renders a
  receipt and why forks can start from an error state.

---

## 6. Fingerprints and comparability

Two runs are **directly comparable** if their fingerprints match. A
fingerprint is a 12-hex-character SHA-256 over the run's
reproducibility inputs, and the rule for what counts as one is
mechanical: every event whose type name ends in `Artifact` feeds the
hash, in the order the run emitted them, canonicalized to JSON and
joined by a record separator. Output events like `ReportFile` never
feed it.

Today that set is `FetchArtifact`, `AnalyzeArtifact`,
`DatasetTasksArtifact`, `SetupArtifact`, `AtomizeArtifact`, and
`ModelProxyArtifact`. Each artifact contributes only its load-bearing
fields, not its whole payload, so per-run noise stays out: a proxy's
hostname and ephemeral port, a sandbox-local file path, which retry
produced the code, the summary the agent wrote about it. There is one
rule per artifact type and one place it is written down, so the hash
recorded when a run finishes is the hash you get back when you ask how
it was derived.

### "Fingerprint" names a family, not one value

The word shows up on several columns, and they answer different
questions. Only the first row is the run fingerprint, which is what
the rest of this section means by "fingerprint".

| Fingerprint | Hashes | Answers | Model in it? |
|---|---|---|---|
| `traces.fingerprint_hash` | the run's `*Artifact` events | Is this the same experiment, subject included? | Yes |
| `benchmark_subset_impls.impl_fingerprint` | `task_impl.py` content | Is this the same scorer code? | No |
| `benchmark_subsets.subset_fingerprint` | the task set (id, input, expected) | Is this the same data? | No |
| `build_subset_revisions.tasks_fingerprint` | the task set as of one edit | What did this edit change? | No |
| `materialize_recommendations.fingerprint_hash` | the artifacts in effect mid-run, latest one per type | What task timeouts and batch size worked last time? | Yes |
| `phase_artifact_cells.artifact_fingerprint` | one phase's output artifact | Can a later run reuse this phase's work? | Only for Atomize |
| `phase_snapshot_pointers.artifact_fingerprint` | the same value as the row above | Which sandbox snapshot holds that artifact? | Inherited |
| `phase_fingerprints.cumulative_hash` | every artifact up to a phase boundary | Do two runs agree with each other up to here? | Only after Attach |

That last column is the whole design. Two different questions get
called "fingerprint", and each one gets its own hash. The run
fingerprint asks *is this the same experiment, subject included*, so
the model belongs in it. The subset, impl and pre-Attach phase
fingerprints ask *is this the same evaluation setup*, so the model is
deliberately kept out of them and results stay comparable across
models. Neither hash is trying to do the other's job.

Attach is the seam. Phases that run before the model is attached
(Fetch, Analyze, Setup) hash to model-free values, which is why a
sweep of ten models over one benchmark downloads and installs it
once and reuses that work nine times. Atomize runs after Attach and
folds the model config into its cache key, because the scorer the
Driver writes can legitimately differ per model.

**The registry never groups by the run fingerprint.** Leaderboards,
run counts and covering-run counts key on
`(benchmark_id, subset_id, impl_id)` with `model` as its own
column (§7). That is why a model sweep still lines up as one
leaderboard even though every run in it carries a distinct
fingerprint.

**The model IS part of the fingerprint.** `ModelProxyArtifact`
contributes exactly two fields, the proxy's model id and its LiteLLM
config YAML, so two runs that differ only in which model they
evaluated get *different* fingerprints. If you want to compare across
models, match on `(benchmark_id, subset_id, impl_id)` (§7) rather
than on the fingerprint. A model sweep produces N runs that share
those three ids and N distinct fingerprints.

This is also why **impls are immutable** (§7). Publishing a new
version of a scorer produces a new impl ID, which produces a new
fingerprint. Old runs keep pointing at the old impl; their
comparability is preserved.

**Trials and comparability.** The trials/aggregator config (§7) is
genuinely *not* part of the fingerprint: it is frozen on the run as
`traces.trials_config`, which is a column, not an artifact. But it
does affect numerical comparability between two runs that match. A
`k=1` run and a `k=5, mean` run on the same fingerprint are measuring
the same evaluation through different lenses, and both numbers are
valid; they just aren't apples-to-apples. When sweeping models you
want to keep `trials` constant across the sweep; when comparing trial
strategies you want to keep the model constant.

---

## 7. Benchmarks, subsets, impls, scorers

| Concept   | What it is                                                                                                                                                                       |
|-----------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| Benchmark | A dataset family (MMLU, BFCL, GSM8K). Metadata only: name, tags, source URL.                                                                                                     |
| Subset    | A split within a benchmark (MMLU "abstract algebra"). Holds the task rows.                                                                                                       |
| Impl      | A scorer (`task_impl.py`) bound to a subset. Immutable.                                                                                                                          |
| Task      | One row from a subset plus the per-run result for it on a specific run.                                                                                                          |
| Trial     | One execution of the scorer against the model for one task. A task with `k=1` has one trial; `k>1` produces `k` trials per task that an aggregator reduces into the per-task score. |

A **scorer** is what Benchwright calls the Python class inside
`task_impl.py`. It implements three methods:

- `extract(row)`: pull the input the model should see out of a raw
  subset row.
- `solve(row, call_sut)`: call the model (via the proxy) and return
  its response.
- `score(row, response)`: compare the response to ground truth and
  return a score in `[0, 1]` plus optional metadata.

The scorer is the only code that can call the model during
Materialize. It runs inside the work sandbox (or a task sandbox in
fanout mode), so it can't see the real provider credentials and
can't reach anything outside the host policy allowlist (§9 of
api.md).

When you "publish a benchmark", you are publishing a
(benchmark, subset, impl) triple. The benchmark and subset can have
many impls over time; each impl is an immutable snapshot.

**Flags that change how a subset behaves:**

- **`official`** (on an impl): marks it as the canonical scorer for
  its subset (admin-promoted). `impl="latest:official"` is the pinned
  launch shape intended to pick the most recent official impl; the
  CLI accepts the flag, but the server does not resolve it yet (§3).
- **`rolling`** (on a subset): the task list is allowed to grow
  over time. New tasks are appended; old tasks are never removed or
  renumbered. Runs against a rolling subset pin to the snapshot the
  run saw at launch, so a run from yesterday stays reproducible
  even after today's append.
- **Synthetic subset**: a subset the Driver creates in-pipeline from
  a procedural script rather than sourcing it from an external
  dataset. Emitted as a `SyntheticSubsetCreated` event by the
  Analyze phase when the operator's scoping phrase needs one (for
  example, "one sample from each MMLU category" produces a
  synthetic subset of 57 tasks).

**Leaderboards.** Every subset has a per-subset leaderboard
(`GET /v1/registry/subsets/{id}/leaderboard`) showing the best
score per model across all public runs against any of its impls.
Only runs with `visibility=public` (§12) appear. A subset on a
private benchmark has no public leaderboard.

**Trials and aggregators.** Materialize runs `k` trials per task,
then reduces them into one per-task score. `k = 1` is the default
and the trial layer is invisible — one call to the model, one score,
one row on the heatmap. `k > 1` is how Benchwright handles `pass@k`,
self-consistency / majority voting, mean-of-runs, and similar
multi-sample protocols without changing the scorer.

The launch body's `trials` field overrides the subset's
`default_trials`, which falls back to `{plan: {k: 1}, aggregator:
{name: "any_pass"}}`. The resolved config is frozen on
`traces.trials_config` so replays and forks reproduce the same
multi-sample shape. Built-in aggregators:

- **`any_pass`** — row score is the best trial's score. With
  `score >= 1.0` as the pass threshold this is exact `pass@k`
  semantics.
- **`all_pass`** — row passes only if every trial passed. Score
  collapses to 0 or 1.
- **`mean`** — arithmetic mean of the per-trial scores. Smoothed
  measurement of a non-deterministic model.
- **`majority_on`** — modal value of `score_meta.<field>` across
  trials wins; the row's score is that winner's score. Implements
  self-consistency / majority voting when the scorer puts the
  parsed answer in `score_meta`.
- **`weighted_mean`** — `sum(score * meta.<weight>) / sum(meta.<weight>)`.
  Useful when trials carry a confidence or token-weight you want
  to fold in.
- **`min_above`** — `any_pass`, but only counts trials whose
  `score_meta.<field>` exceeds a threshold. "Row passes if any
  high-confidence trial passed."

A scorer makes per-trial fields available by returning a dict from
`score()` instead of a bare float; the dict's `score` key is the
canonical score and the rest goes into per-trial **`score_meta`**
that the aggregators above can read.

Per-trial events (`MaterializeTaskResult`) are stamped with a
0-based `trial` index and the same `(invocation, index)`. After all
trials for a task settle, one `MaterializeTaskAggregated` event
publishes the row-level score the dashboard and leaderboard read.
The dashboard's modal lets you drill from the aggregated row down
into the per-trial scores and outputs.

### Run shape: what makes one run distinct from another

A run's identity is a **4-tuple**: `(benchmark_id, subset_id, impl_id, model)`.

| Dimension     | What it means                                                   | Where it lives                                       |
|---------------|------------------------------------------------------------------|------------------------------------------------------|
| `benchmark_id`| Which benchmark family (MMLU, BFCL, …)                          | `benchmark_run_summaries.benchmark_id`               |
| `subset_id`   | Which dataset bytes (split + selection + content hash)          | `benchmark_run_summaries.subset_id`                  |
| `impl_id`     | Which task_impl.py code                                          | `benchmark_run_summaries.impl_id`                    |

For a benchwright-BUILT (product) benchmark the `task_impl.py` is generated:
it is the product runner plus the subset's **recipe** header, so `impl_id` is a
hash of the recipe. Editing the recipe mints a new impl; you do not register or
select one directly the way you can for a hand-written scorer.
| `model`         | Which model was attached                                         | `traces.model` (jsonb: `{source, identifier}`)         |

These dimensions are **deliberately orthogonal**. In particular:

- **`impl_id` is code-only, scoped to its benchmark.** It's
  `<benchmark_id>:<first 12 hex of sha256(task_impl.py content)>`;
  the full digest is kept on `benchmark_subset_impls.impl_fingerprint`.
  (Benchwright-built benchmarks key the same shape on the subset id
  instead of the benchmark id.) Two byte-identical scorer files share
  an `impl_id` regardless of which model they ran against, but only
  within the same benchmark: the same scorer under two benchmarks
  gets two ids. This lets you ask "is this the same code as run X"
  without first matching on the model.
- **model is a separate axis, not folded into `impl_id`.** Even though
  the Atomize *cache key* includes `model_config_hash` (so swapping the
  model typically misses cache and triggers a fresh `task_impl.py`),
  the resulting impl is identified by content alone. If two models
  happen to produce identical code, they share an impl.

**What the ids are good for:**

- The **identicon** (the colored tile next to a benchmark name on the
  registry's benchmark cards and detail header) is a deterministic
  tile derived from the benchmark id alone. Same benchmark, same
  tile, on every card and every list that shows it. It is a registry
  affordance, not a run one: the run dashboard does not render it.
- The full **4-tuple** is what defines "the same run shape". It's
  what replay/dup detection and reproducibility comparisons use:
  to claim "this is a true replay of run X", all four must match.
  A rerun against a new model shares the first three but not the
  fourth — that's a different run, just one that's deeply related.

The `benchmark_run_summaries` projection populates `benchmark_id` /
`subset_id` / `impl_id` from the trace events on terminal status
(via a Postgres trigger). In-flight runs return `null` for those
columns, so anything keyed on the run shape only resolves once the
run reaches `ok` / `error` and the trigger fires.

---

## 8. Driver vs SUT vs Operator

Three different LLM-shaped things, not interchangeable, and
conflating them is the most common source of confused bug reports.

- **Driver**: the locked-down LLM that executes one run (defaults to
  `openrouter://deepseek/deepseek-v4-flash`). Its whole capability set is monitor,
  debug, and re-run a phase: it reads phases, writes `task_impl.py`,
  and decides when a phase is done. It is scoped to its run and
  cannot launch another one. You pay for Driver tokens out of your
  balance.
- **SUT**, the system under test: the thing being evaluated. Usually
  an LLM named `provider/model` (`openrouter/qwen/qwen3-32b`), but
  it could be any HTTP endpoint the proxy can front. This is what
  `run.model` names, and the score is a statement about it.
- **Operator**: the unrestricted, goal-driven agent that works on
  your behalf across runs. It launches runs, picks each one's
  Driver, reads the metrics that come back, and re-runs with a
  different Driver when one struggles. The Driver re-runs *phases*;
  the Operator re-runs *runs*.

"Model" on its own is ambiguous. In API fields, `run.model` is
always the SUT. In casual speech, "the model" often means the
Driver. If a doc or log line says "model" without context, assume
SUT.

The v1 API used to blur exactly this line: the response fields
`run.operator` and `cost.operator` predated the split and carried the
**Driver's** model id and the Driver's LLM spend, not the Operator's.
Both were removed. `run.driver` and `cost.driver` carry those two
values under the right name and are the only keys for them. The SUT's
own spend is `cost.model`.

Input matches. The launch body key is `driver` and nothing else; the
CLI flag is `--driver`; the Python SDK kwarg is `driver=`. The retired
`operator` and `agent` spellings are rejected on every one of those
surfaces rather than quietly accepted, so a caller still using one
finds out instead of running on an unintended model.

Events mark the distinction where it matters: `AgentTextEvent` is
the Driver talking, `MaterializeTaskResult` is about the SUT.

---

## 9. Cost, constraints, and halting

Cost is reported under `run.cost`, with each source on its own key so
you never have to guess which LLM a number refers to:

- **`cost.driver`**: tokens the **Driver** spent (§8). The only key
  for it; `cost.operator` and `cost.agent` were removed.
- **`cost.model`**: tokens the **SUT** spent during Materialize.
  Separate key, not folded into the Driver's.
- **`cost.sandbox`**: the sum of `(duration, instance-size)` across
  all sandboxes in the run, charged at the backend's rate (sandbox0
  or E2B).
- **`cost.narration`**: the website's run-summary LLM spend, with
  call and token counts beside it.
- **`cost.billed`**: what actually hits your balance, which is the
  platform portion (sandbox, Driver, narration) with markup applied.

`run.cost.incurred` is the sum of those sources. It is updated on every
phase tick, not just at the end, so the billing panel and any
external ledger can track burn in real time.

One exception: **`cost.proxy_linger` is not in `total`**. The
post-run proxy linger window is metered after the run's own ledger is
finalized, so it is reported separately, and it carries an exact
`amount_usd` alongside the cents because a linger is routinely
sub-cent. Add it at the read site if you want an all-in figure.

**Limits.** A launch declares its ceilings in a `limits` block, nested
`<level>.<quantity>`, with money in whole **dollars** and time in
**seconds** — the unit belongs to the quantity, never to the name:

```json
{"limits": {"run": {"cost": 5.0}, "task": {"attempts": 1},
            "attempt": {"duration": 600}}}
```

Levels are `run`, `task`, `attempt` and `phase` (which may name one,
`phase:atomize`). A `run` or `phase` ceiling halts the run with
`halt_reason=limit_exceeded:<level>.<quantity>`; a `task` ceiling fails
only that task, never retries it, and lets the run continue. A limit the
platform cannot enforce is **refused with 400** rather than stored and
ignored. The full matrix is in the [CLI reference](./cli.md#limits).

`run.cost` compares `cost.incurred`, which includes model spend — so it
fires on a BYOK run where the billed figure would barely move.

**Constraints (older spelling).** `constraints` still works and means the
same thing, but carries units in its key names:

- `max_cost_cents`: `run.cost`, in cents rather than dollars.
- `max_duration_ms`: `run.duration`, in milliseconds rather than seconds.
- `max_tasks`: **caps the sample count**; it does not halt. It is
  folded together with `metadata.limit` and the harness config's own
  `limit`, and the narrowest of the three wins. A run that hits the
  cap finishes normally over the smaller sample, with `status = ok`.

The two halting constraints are checked by the same live loop that
tracks balance, and they trip the moment they're exceeded rather than
trimming work to stay under. `max_tasks` is the opposite: it is a
scoping choice made at launch, applied before any task runs.

**Halt mechanism.** Halts flow through one mechanism. When a halt
condition trips (balance exhausted, constraint exceeded, manual stop,
admin intervention), the server flips `halted = true` and sets a
`halt_reason`. The pipeline polls that flag every couple of seconds
between phases and tears itself down cleanly: proxy and work
sandboxes get killed, and partial results are finalized. `halted` is
a boolean beside `status`, not a status of its own (§1); read
`display_status` if you want one field that says "halted".

The `halt_reason` values in use:

| Reason | Meaning |
|---|---|
| `insufficient_balance` | Balance ran out mid-run. |
| `token_cap` | The launching PAT hit its spend cap. |
| `limit_exceeded:<level>.<quantity>` | A `limits` ceiling tripped, e.g. `limit_exceeded:run.cost`. The suffix says which. |
| `max_cost_exceeded` | Pre-`limits` spelling of `limit_exceeded:run.cost`. |
| `max_model_cost_exceeded` | Pre-`limits` spelling of `limit_exceeded:run.model_cost`. |
| `max_duration_exceeded` | Pre-`limits` spelling of `limit_exceeded:run.duration`. |
| `compute_constraints` | No compute could be allocated for the run. |
| `sandbox_died` | A sandbox the run depended on went away. |
| `spawn_failed` | A task sandbox could not be created. |
| `checkpoint_relaunch` | Deliberate pause between relaunch legs. |
| `agent_review: <reason>` | The Driver escalated for human review (§10). |
| `restart_cap_exceeded: <reason>` | A phase restarted too many times (§10). |
| `operator_pause: <principal>` | Someone paused the run. |
| `operator_kill by <principal>` | Someone killed the run. |
| `admin_halt by <admin>` | An admin stopped the run. |

Several of these carry a suffix, so **match on prefix, not on
equality**. `halt_reason.startsWith("operator_kill")` is right;
`== "operator_kill"` misses every real one.

---

## 10. Forks, replay, and escalation

Three ways the pipeline can re-enter a phase. Same concept at three
scopes.

**Replay** re-runs a reference run with a different model. The
benchmark, subset, impl, and sample selection are all pinned to the
reference; only the model changes. Use replay for noise checks
(same model, same run, just rerun), A/B tests, and cross-model
sweeps.

**Fork** rewinds to a specific phase in an existing run and
continues from there with overrides. Common uses:

- Rewind to Atomize, rewrite the scorer, continue to Materialize.
- Rewind to Attach, swap the model, continue without re-fetching or
  re-analyzing.
- Rewind to Fetch after the original source URL broke, point at a
  mirror, continue.

A fork inherits the parent's event log up to the rewind point, then
its own log picks up from there. The parent's artifacts and work
sandbox snapshots are reused; nothing upstream of the rewind point
is recomputed. This is how forks are cheap: you pay only for the
phases that actually run.

Replay is a special case of fork that rewinds all the way to
Materialize (everything above is pinned) and swaps the model.

**Phase restart** is the in-run equivalent: the Driver itself can
request the current phase start over, usually because a partial
result was bad. Restarts increment the phase's invocation counter
(§2) but stay inside the same run; they don't create a new run ID
and they don't touch earlier phases. Atomize restarts are the most
common case: the Driver writes a scorer, runs it on sample tasks,
decides the scorer is wrong, and requests a restart to rewrite it.

Restarts are capped at two per run. A phase asking for a third
auto-escalates to the same waiting state as human review, with
`halt_reason = restart_cap_exceeded: <reason>`. This is the loop
guard: a Driver that can't converge parks the run for a human
instead of burning your balance on ping-pong.

**Human review** is the escalation path when the Driver can't make
progress without a human decision. Triggered by the Driver via the
`requestHumanReview` tool, it halts the run with
`halt_reason = agent_review: <reason>` and waits, showing as
`display_status = waiting`.

Resolve it with `PATCH /v1/runs/{id}` carrying a `review_resolution`
string, which is your guidance in prose. That write clears
`halted` / `halt_reason` in the same row update and wakes the driver
sandbox, so the run picks up with your text as its instructions.

`POST /v1/runs/{id}/resume` is a different door: it refuses anything
that isn't an `operator_pause` halt and returns `409 not_paused`, so
it is not the way out of a review. Three halt reasons are resumable
at all: `agent_review`, `operator_pause`, and
`restart_cap_exceeded`. A balance halt is not among them.

---

## 11. Host policy and network isolation

Scorers run untrusted. Your impl could do `requests.get(...)` on any
URL, and the pipeline has to assume it will. Two layers keep that
safe.

**The proxy sandbox (§4)** is the first layer. The work sandbox
does not have the real model API key. It has a fake key pointed at
the proxy, and the proxy holds the real credentials. A scorer that
tries to call the model directly (bypassing the proxy) gets no
credentials and no route.

**Host policy** is the second layer. Every outbound request from
the work sandbox is filtered against your account's host policy.
A policy entry tags a host pattern as either `allow` (permitted)
or `block` (refused). There is a default mode that applies to
anything not listed.

- **Default mode** is `allow` on a fresh account, which is
  permissive but makes first-run ergonomics easy. Flip to `block`
  to make the allowlist exhaustive.
- **Allow rules** explicitly permit a host. Adding a provider
  secret with a `host_pattern` auto-upserts an allow rule so your
  proxy stops blocking it.
- **Block rules** explicitly refuse a host even when default mode
  is `allow`. Use this for hosts you want to ban regardless of how
  permissive the default is.

**Blocked-host log.** Every refusal is recorded with a
`seen_count`, `first_seen`, `last_seen`. The workflow is: run
once, read the blocked-host log (`GET /v1/runs/{id}/blocked-hosts`
or the aggregated `GET /v1/blocked-hosts`), decide which to allow,
rerun. This makes "trust on observation" explicit: you never have
to guess which hosts a benchmark actually needs.

---

## 12. Task outcomes and visibility

**Task outcomes.** Every settled task row on a run lands in one of
three buckets, distinguished on the dashboard heatmap by color. The
API values are `pass`, `fail`, and `err`; the dashboard labels them
"Passed", "Failed", and "Errored". A task still in flight reports a
fourth value, `running` (teal), until it settles.

- **`pass`** (green): the scorer returned `score >= 1.0`. The model
  answered correctly.
- **`fail`** (yellow): the scorer ran cleanly and returned
  `score < 1.0`. The model answered, but the answer was wrong or
  partial. This is the expected outcome for any task the model
  doesn't know; it is *not* a failure of the pipeline.
- **`err`** (red): the scorer couldn't produce a score at all
  (exception during `extract`/`solve`/`score`, proxy timeout,
  sandbox crash). Errors are pipeline/infra failures; failures are
  model failures. Keep them separate when investigating.

When the run has `k > 1` trials per task (§7), the bucket is
determined by the *aggregated* score, not the per-trial scores. A
task with five trials at `0.0, 0.0, 1.0, 0.0, 1.0` and an
`any_pass` aggregator lands in `pass`; the same trials with
`all_pass` land in `fail`. The dashboard exposes the per-trial
breakdown on click, but the heatmap colour is always the
post-aggregation outcome.

**Pass rate** and **accuracy** are the same projection:
`passed / (passed + failed)`. Errors are excluded from the
denominator, so a run with many errors can still report a high
pass rate over the tasks that did complete. The dashboard uses
"accuracy" for the hero tile and "pass rate" everywhere else;
treat the terms as synonyms.

**Run visibility.** Every run has a visibility setting controlling
who can see it:

- **Private** (default): only the run's owner can load it. The
  run ID is not enumerable by anyone else.
- **Unlisted**: anyone with the URL can load the run page, but the
  run is not indexed or listed in registry/search results.
  Useful for sharing a result with a colleague without publishing.
- **Public**: the run is listed on its parent benchmark's registry
  page and appears in leaderboards. Requires the parent benchmark
  to also be public; you cannot publicly list a run against a
  private benchmark.

Visibility is per-run, not per-account: you can keep most runs
private and flip individual runs to unlisted or public. Changing
visibility is reversible at any time.

---

## 13. Industry terminology crosswalk

If you've used HumanEval, BFCL, lm-evaluation-harness, HELM, or
similar, the terms you know translate like this:

| Industry term                       | Benchwright equivalent                                  |
|-------------------------------------|----------------------------------------------------------|
| Eval / evaluation                   | Run (§1)                                                 |
| Eval harness                        | Scorer + pipeline. The scorer is `task_impl.py` (§7); the pipeline runs it. |
| Eval orchestrator / runner agent    | Driver (§8). One locked-down LLM scoped to one run. It re-runs *phases* inside its run; it never launches another run. |
| (no common term)                    | Operator (§8). The unrestricted goal agent above the runs: it launches them, picks each one's Driver, reads the metrics, and re-runs with a different Driver. The Driver re-runs phases; the Operator re-runs runs. |
| Benchmark / dataset                 | Benchmark (§7). The dataset lives under its subsets.    |
| Split (train / dev / test)          | Subset (§7). There is no train/dev/test distinction: every Benchwright subset is test-only. |
| Task / example / problem / instance | Task (§7). One row of a subset plus the per-run result. |
| Prompt / input                      | Whatever `scorer.extract(row)` returns (§7).            |
| Completion / generation / response  | What `scorer.solve(row, call_sut)` returns.             |
| Ground truth / gold / reference     | The `expected` field on a task row.                     |
| Zero-shot / k-shot / few-shot       | Scorer's choice inside `extract`. The pipeline does not enforce or count in-context examples. |
| Chain-of-thought (CoT)              | Scorer pattern, not a pipeline knob. Scorers can request or parse reasoning in their prompt/response. |
| `pass@k`                            | `trials.plan.k = k` with `aggregator = any_pass` (§7). Each task gets `k` trials; the row passes if any one does. |
| Self-consistency / majority voting  | `trials.plan.k > 1` with `aggregator = majority_on`, where the scorer puts the parsed answer in `score_meta` (§7). |
| Best-of-N / `best@k`                | Same as `pass@k` for binary scoring. For graded scoring, `any_pass` returns the highest-scoring trial. |
| Multi-sample evaluation             | Trials. One trial = one sample (§7).                     |
| Stochastic / non-deterministic eval | Use `trials.plan.k > 1` with `aggregator = mean` to smooth temperature-driven variance. |
| Attempt / try / sample              | Trial (§7). The 0-based `trial` index in `MaterializeTaskResult`. |
| Exact match / F1 / BLEU / ROUGE     | All normalized to a single `[0, 1]` score (§12). The method name is stored on the *benchmark* as `eval_method` for documentation; the pipeline sees only the float. |
| Metric                              | The `[0, 1]` score. Benchwright is single-metric by design; secondary metrics go in the per-task `meta` JSON alongside the score. |
| Pass rate                           | Pass rate (§12). `passed / (passed + failed)`; errors excluded from the denominator. |
| Accuracy                            | Synonym for pass rate (§12).                             |
| Leaderboard                         | Leaderboard (§7). First-class, per-subset, ranked by best score per model across public runs. |
| Contamination / test leakage        | Not tracked. Operator's responsibility to verify a subset is clean before publishing. |
| Hyperparameters (temperature, top_p, max_tokens) | Passed through to the SUT via the proxy sandbox. Set once at launch; the pipeline does not vary them per-task. |
| Context window / context length     | A property of the SUT, not a Benchwright knob. The scorer is responsible for staying within it. |
| Token / tokens                      | Counted per LLM role: `cost.driver` for Driver tokens, `cost.model` for SUT tokens (§9). Tokenization itself is done by the model or proxy, not the pipeline. |
| Golden dataset / eval set           | A published subset (§7). Make it `rolling` if it grows.  |
| Benchmark suite / multi-task bench  | Not implemented. Score each benchmark as its own run and compare across them; there is no subset that spans benchmarks. |

**Where Benchwright differs from convention:**

- **`pass@k` and friends are a config knob, not a scorer
  responsibility.** Set `trials.plan.k` at launch and pick an
  aggregator (§7). The scorer keeps doing one call to the model and
  one score; the pipeline runs `k` independent trials and reduces
  them server-side. The same scorer code supports `pass@k`,
  self-consistency, mean-of-N, and weighted-mean without
  modification.
- **No explicit splits.** If a source dataset ships a train/test
  split, the Driver discovers each as its own subset during
  Analyze. There is no "this is the training half, don't eval on
  it" semantics; you simply don't publish a subset you don't want
  people evaluating on.
- **No per-task hyperparameters.** SUT sampling params are set
  once on the proxy at Attach time. A scorer that needs different
  temperatures per task must route through multiple proxy
  configurations, which today means multiple runs.
- **One score, one number.** Benchwright deliberately collapses
  every metric to `[0, 1]`. This is what makes leaderboards and
  replays comparable across scorers. Scorers that want to expose
  secondary measurements (BLEU, latency, tokens-used, etc.) put
  them in the per-task `meta` JSON, which the dashboard can
  display and the API returns, but which does not affect ranking.
