Public benchmarks

Benchwright Registry

Open benchmark registry. Browse, run, and compare models — every public run lands here for everyone.

benchmarks
subsets
tasks
total runs
models
⌘K
back to registry
Built with the benchmark builder — 1 task · llm-judge grader. See how it was built →

Pokey bench (v2: lean state)

Pallet Town starter + reach Viridian; fixed dialogue + seamless walk/transition + lean state payload (drops redundant collision grid) to keep SUT context small.

1 sample 1 run · 1 impl impl-ffad39047a55
metricscore
What it measures
othertextmultiple-choice
Pallet Town starter + reach Viridian; fixed dialogue + seamless walk/transition + lean state payload (drops redundant collision grid) to keep SUT context small.
How it's evaluated
Grader
impl-ffad39047a55
Task count
1
Category
other
Avg duration
43m 59s
Avg tokens
4.4M in → 16.6k out
Platform cost
$0.1423
Platform cost is what Benchwright charges to run the eval (sandbox execution) — roughly model-independent. Model inference cost is paid to your provider and depends entirely on the model you pick (e.g. an Opus run costs far more than a small open model); see the leaderboard for per-model cost.
Top models
1
gpt-6-astra
43m 59s
1.000
Run it yourself
Pick the official fingerprint, give it a model, and benchwright will reproduce the exact evaluation. Public runs land back in the registry.
Download dataset (JSONL)1 task · 395 B