Public benchmarks

Benchwright Registry

Open benchmark registry. Browse, run, and compare models — every public run lands here for everyone.

benchmarks
subsets
tasks
total runs
models
⌘K
back to registry
Built with the benchmark builder — 1 task · llm-judge grader. See how it was built →

Starter & Viridian (agent, colour, 0.5s)

Long-form Pallet Town on NousResearch/pokemon-agent in Game Boy COLOR mode. Frames sampled at 0.5s and encoded to a film. Leave the house, get a starter, reach Viridian City; staged partial credit.

1 sample 1 run · 1 impl impl-84d7fc2f94c1
metricscore
What it measures
agenticreward
Long-form Pallet Town on NousResearch/pokemon-agent in Game Boy COLOR mode. Frames sampled at 0.5s and encoded to a film. Leave the house, get a starter, reach Viridian City; staged partial credit.
How it's evaluated
Grader
impl-84d7fc2f94c1
Task count
1
Category
agentic
Avg duration
27m 31s
Avg tokens
9.5M in → 21.1k out
Platform cost
$0.0185
Platform cost is what Benchwright charges to run the eval (sandbox execution) — roughly model-independent. Model inference cost is paid to your provider and depends entirely on the model you pick (e.g. an Opus run costs far more than a small open model); see the leaderboard for per-model cost.
Top models
1
google/gemini-3.7-flash
27m 31s
1.000
Run it yourself
Pick the official fingerprint, give it a model, and benchwright will reproduce the exact evaluation. Public runs land back in the registry.
Download dataset (JSONL)1 task · 1 KB