Public benchmarks

Benchwright Registry

Open benchmark registry. Browse, run, and compare models — every public run lands here for everyone.

benchmarks
subsets
tasks
total runs
models
⌘K
back to registry

Aime2026

inspect harness benchmark (inspect_evals/aime2026)

source 30 samples 1 run · 1 impl impl-11b01bc914e3
What it measures
mathtextexact-match
inspect harness benchmark (inspect_evals/aime2026)
How it's evaluated
Grader
impl-11b01bc914e3
Task count
30
Category
math
Avg duration
10m 4s
Avg tokens
6.7k in → 363.9k out
Platform cost
$0.1782
Platform cost is what Benchwright charges to run the eval (sandbox execution) — roughly model-independent. Model inference cost is paid to your provider and depends entirely on the model you pick (e.g. an Opus run costs far more than a small open model); see the leaderboard for per-model cost.
Official metricsevery metric the harness exported
accuracy
best 0.767
accuracy_stderrstandard error
best 0.077
Top models
1
openai/gpt-5-nano
10m 4s
0.767
Run it yourself
Pick the official fingerprint, give it a model, and benchwright will reproduce the exact evaluation. Public runs land back in the registry.
Download dataset (JSONL)30 tasks · 13 KB