Platform cost is what Benchwright charges to run the eval (sandbox execution) — roughly model-independent. Model inference cost is paid to your provider and depends entirely on the model you pick (e.g. an Opus run costs far more than a small open model); see the leaderboard for per-model cost.
Official metricsevery metric the harness exported
accuracy
best 0.767
accuracy_stderrstandard error
best 0.077
Top models
1
openai/gpt-5-nano
10m 4s
0.767
Run it yourself
Pick the official fingerprint, give it a model, and benchwright will reproduce the exact evaluation. Public runs land back in the registry.