Public benchmarks

Benchwright Registry

Open benchmark registry. Browse, run, and compare models — every public run lands here for everyone.

benchmarks
subsets
tasks
total runs
models
⌘K
back to registry

runebench

An agentic terminal/CLI benchmark where each task is a containerized environment testing an agent's ability to operate a shell to solve real-world problems.

1 sample 1 run · 1 impl impl-c846f58a6753
metricscore
What it measures
agenticreward
An agentic terminal/CLI benchmark where each task is a containerized environment testing an agent's ability to operate a shell to solve real-world problems.
How it's evaluated
Grader
impl-c846f58a6753
Task count
1
Category
agentic
Avg duration
15m 40s
Avg tokens
1.7M in → 14.1k out
Platform cost
$0.0184
Platform cost is what Benchwright charges to run the eval (sandbox execution) — roughly model-independent. Model inference cost is paid to your provider and depends entirely on the model you pick (e.g. an Opus run costs far more than a small open model); see the leaderboard for per-model cost.
Top models
1
deepseek/deepseek-v4-pro
15m 40s
28
Run it yourself
Pick the official fingerprint, give it a model, and benchwright will reproduce the exact evaluation. Public runs land back in the registry.
Download dataset (JSONL)1 task · 2 KB