⌘K
Open benchmark registry. Browse, run, and compare models — every public run lands here for everyone.
An agentic terminal/CLI benchmark where each task is a containerized environment testing an agent's ability to operate a shell to solve real-world problems.