Coding · benchmark
Terminal-Bench
Give agents executable tasks in real terminal environments.
What it measures
Task resolution rate (%)
Pin the benchmark release, agent harness and execution budget.
Public evaluation code. Check upstream code and dataset terms.
Open source / dataset · Run instructions · Official results
Reported ranking tracks
No numeric ranking has been imported for this project. Read the upstream results and instructions.