Coding · benchmark
SWE-bench-Live
Evaluate coding agents on a continually refreshed collection of repository issues.
What it measures
Resolved tasks (%)
Use the same dated release and submission rules when comparing runs.
Public evaluation code. Check upstream code and dataset terms.
Open source / dataset · Run instructions · Official results
Reported ranking tracks
No numeric ranking has been imported for this project. Read the upstream results and instructions.