Coding · benchmark
SWE-bench Verified
Fix real GitHub issues and check patches with repository tests.
Resolved tasks (%)
Open source / dataset · Official resultsBUILDERWARS EVALS
50 benchmarks and frameworks. Find the right test, read the reported rankings, then build a better contender.
Explore reported rankings · Find a fun competition
Coding · benchmark
Fix real GitHub issues and check patches with repository tests.
Resolved tasks (%)
Open source / dataset · Official resultsCoding · benchmark
Evaluate coding agents on a continually refreshed collection of repository issues.
Resolved tasks (%)
Open source / dataset · Official resultsCoding · benchmark
Test code generation on programming problems with explicit time windows.
Pass@1 (%)
Open source / dataset · Official resultsCoding · benchmark
Stress-test HumanEval and MBPP solutions with expanded test cases.
HumanEval+ / MBPP+ pass@1 (%)
Open source / dataset · Official resultsCoding · benchmark
Measure practical Python generation involving libraries and complex instructions.
Pass@1 (%)
Open source / dataset · Official resultsCoding · benchmark
Give agents executable tasks in real terminal environments.
Task resolution rate (%)
Open source / dataset · Official resultsCoding · benchmark
Evaluate scientific code generation with domain knowledge and execution tests.
Solved scientific problems
Open source / dataset · Official resultsCoding · benchmark
Compare coding models through Aider across multiple programming languages.
Exercises solved (%)
Open source / dataset · Official resultsTool use · benchmark
Check single-turn, multi-turn and agentic function calls.
BFCL V4 overall accuracy (%)
Open source / dataset · Official resultsTool use · benchmark
Evaluate customer-service agents on tool use, policy compliance and user interaction.
Pass^1 and reliability
Open source / dataset · Official resultsTool use · benchmark
Test stateful tool use with intermediate milestones and conversational dependencies.
Milestone-based task success
Open source / datasetWeb agents · benchmark
Complete realistic tasks on reproducible, self-hosted websites.
Task success rate (%)
Open source / dataset · Official resultsWeb agents · benchmark
Solve web tasks requiring screenshots and visual reasoning.
Task success rate (%)
Open source / dataset · Official resultsWeb agents · benchmark
Measure enterprise workflow automation in browser-based environments.
Task success rate (%)
Open source / dataset · Official resultsWeb agents · benchmark
Evaluate grounding and action prediction across diverse websites.
Element / action accuracy
Open source / dataset · Official resultsComputer use · benchmark
Test computer-use agents on real desktop applications and files.
Task success rate (%)
Open source / dataset · Official resultsGeneral agents · benchmark
Challenge general assistants with research, tool use and multimodal questions.
Answer accuracy by difficulty
Open source / dataset · Official resultsReasoning · benchmark
Evaluate broad knowledge and reasoning with more challenging multiple-choice questions.
Overall accuracy (%)
Open source / dataset · Official resultsReasoning · benchmark
Measure knowledge across academic and professional subjects.
Subject and aggregate accuracy
Open source / datasetReasoning · benchmark
Test graduate-level science reasoning on expert-authored questions.
Accuracy (%)
Open source / datasetReasoning · benchmark
Evaluate difficult expert questions spanning science, mathematics and humanities.
Accuracy and calibration error
Open source / dataset · Official resultsReasoning · benchmark
Probe reasoning through challenging tasks selected from BIG-Bench.
Task accuracy
Open source / datasetReasoning · benchmark
Solve novel grid transformations with strict separation between training and evaluation.
Solved tasks and cost per task
Open source / dataset · Official resultsGames · benchmark
Evaluate adaptation in novel interactive environments.
Interactive task performance
Open source / dataset · Official resultsMathematics · benchmark
Test multi-step grade-school mathematical word problems.
Answer accuracy
Open source / datasetMathematics · benchmark
Evaluate competition-level mathematical problem solving.
Exact-answer accuracy
Open source / datasetMathematics · benchmark
Compare reasoning models on dated mathematical competitions.
Accuracy by competition
Open source / dataset · Official resultsMultimodal · benchmark
Test expert-level reasoning across images and academic disciplines.
Validation / test accuracy (%)
Open source / dataset · Official resultsMultimodal · benchmark
Use a harder visual reasoning track with stronger evaluation controls.
Pro overall / vision accuracy (%)
Open source / dataset · Official resultsMultimodal · benchmark
Combine visual perception and mathematical reasoning.
Answer accuracy (%)
Open source / dataset · Official resultsLong context · benchmark
Evaluate long-context understanding across tasks and languages.
Task-specific score / V2 accuracy
Open source / dataset · Official resultsLong context · benchmark
Measure effective context length with controlled synthetic tasks.
Accuracy by context length
Open source / datasetInstruction following · benchmark
Check instruction-following constraints with deterministic verifiers.
Strict / loose instruction accuracy
Open source / datasetSafety · benchmark
Evaluate safety behavior against standardized harmful-behavior test cases.
Attack success / refusal behavior
Open source / dataset · Official resultsSafety · benchmark
Measure resistance to common falsehoods and misconceptions.
Truthfulness / multiple-choice score
Open source / datasetGames · benchmark
Evaluate and train language-model agents in competitive and cooperative text games.
Game-specific outcomes and ratings
Open source / dataset · Official resultsEval frameworks · framework
Run reproducible language-model evaluations across a large task library.
Task-defined metrics
Open source / datasetEval frameworks · framework
Build agent, tool-use and model evaluations with reusable scoring and sandboxes.
Evaluation-defined scores
Open source / datasetEval frameworks · framework
Evaluate models using configurable tasks and multiple inference backends.
Task-defined metrics
Open source / datasetEval frameworks · framework
Manage broad language-model evaluation with reusable model and dataset configurations.
Dataset-specific scores
Open source / datasetEval frameworks · framework
Compare language models using explicit scenarios and multiple evaluation dimensions.
Scenario-specific metrics
Open source / dataset · Official resultsEval frameworks · framework
Run multimodal benchmarks through a common evaluation toolkit.
Benchmark-specific metrics
Open source / datasetEval frameworks · framework
Evaluate multimodal models across image, video and audio tasks.
Benchmark-specific metrics
Open source / datasetWeb agents · framework
Run browser-agent tasks through a shared Gym-style environment.
Task success and environment metrics
Open source / datasetEval frameworks · framework
Run isolated agent evaluations and retain trial artifacts across benchmark environments.
Task-defined metrics
Open source / datasetRAG & applications · framework
Evaluate retrieval, faithfulness and answer quality in RAG applications.
Retrieval and generation metrics
Open source / datasetRAG & applications · framework
Compare prompts, applications and agents with assertions and regression tests.
Assertion pass rate / configured scores
Open source / datasetRAG & applications · framework
Test LLM applications with configurable metrics and evaluation datasets.
Metric-specific scores
Open source / datasetSafety · framework
Probe model weaknesses with reproducible scanners and detectors.
Probe / detector outcomes
Open source / datasetGames · framework
Study competitive and cooperative agents across game environments and algorithms.
Game-specific outcomes
Open source / datasetDirectory reviewed 2026-10-07. Dataset access and licenses vary by project.