Public benchmarks

Benchwright Registry

Open benchmark registry. Browse, run, and compare models — every public run lands here for everyone.

benchmarks
subsets
tasks
total runs
models
⌘K
back to registry
Built with the benchmark builder — 1 task · llm-judge grader. See how it was built →

pallet-town-reach-viridian

Harder Pallet Town screen-game task: leave the house, get a starter from Oak, and travel north through Route 1 to reach Viridian City (a strict superset of reach-oaks-lab).

1 sample 1 run · 1 impl impl-0eb0c027779c
metricscore
What it measures
othertextmultiple-choice
Harder Pallet Town screen-game task: leave the house, get a starter from Oak, and travel north through Route 1 to reach Viridian City (a strict superset of reach-oaks-lab).
How it's evaluated
Grader
impl-0eb0c027779c
Task count
1
Category
other
Avg duration
19m 37s
Avg tokens
244.4k in → 60.2k out
Platform cost
$0.0167
Platform cost is what Benchwright charges to run the eval (sandbox execution) — roughly model-independent. Model inference cost is paid to your provider and depends entirely on the model you pick (e.g. an Opus run costs far more than a small open model); see the leaderboard for per-model cost.
Top models
1
google/gemini-3.7-flash
19m 37s
1.000
Run it yourself
Pick the official fingerprint, give it a model, and benchwright will reproduce the exact evaluation. Public runs land back in the registry.
Download dataset (JSONL)1 task · 754 B