REPORTED RANKINGS
Know the score. Know the conditions.
14 ranking tracks. Captured 2026-10-07T07:01:28Z. Scores are reported by the linked benchmark organizers. BuilderWars republishes dated source extracts and does not independently run or certify these evaluations.
SWE-bench Verified · agent systems
500 Verified tasks; Different submitted agent systems and model configurations; this is a system ranking.
All official Verified submissions with a reported score
| Position | Model / system | Resolved tasks (%) |
|---|---|---|
| 1 | live-SWE-agent + Claude 4.5 Opus medium (20251101)live-SWE-agent · Submission · Reasoning medium · System: Attempts - 1 | 79.2% |
| 1 | Sonar Foundation Agent + Claude 4.5 OpusSonar Foundation Agent · Submission · System: Attempts - 1 | 79.2% |
| 3 | TRAE + Doubao-Seed-CodeTRAE · Submission · System: Attempts - 2+ | 78.8% |
| 4 | live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)live-SWE-agent · Submission · System: Attempts - 1 | 77.4% |
| 5 | Atlassian Rovo Dev (2025-09-02)Atlassian Rovo Dev · Submission · System: Attempts - 2 | 76.8% |
Official score source · Retained source extract · Explore this track
SWE-bench Verified · mini-SWE-agent
500 Verified tasks; mini-SWE-agent submissions only, including declared agent versions and settings.
mini-SWE-agent subset of official Verified data
| Position | Model / system | Resolved tasks (%) |
|---|---|---|
| 1 | Claude 4.5 Opus (high)mini-SWE-agent · Submission · Harness version 2.0.0 · Reasoning high · System: Attempts - 1 | 76.8% |
| 2 | Gemini 3 Flash (high)mini-SWE-agent · Submission · Harness version 2.0.0 · Reasoning high · System: Attempts - 1 | 75.8% |
| 2 | MiniMax M2.5 (high)mini-SWE-agent · Submission · Harness version 2.0.0 · Reasoning high · System: Attempts - 1 | 75.8% |
| 4 | Claude 4.6 Opusmini-SWE-agent · Submission · Harness version 2.0.0 · System: Attempts - 1 | 75.6% |
| 5 | Claude 4.5 Opus (20251101) (medium)mini-SWE-agent · Organizer-checked · Harness version 1.16.0 · Reasoning medium · System: Attempts - 1 | 74.4% |
Official score source · Retained source extract · Explore this track
BFCL V4 · overall
V4 overall score; FC and Prompt calling modes remain in model names. Estimated benchmark cost/latency are not independent billing evidence.
All source rows with this metric
| Position | Model / system | Overall accuracy (%) |
|---|---|---|
| 1 | Claude-Opus-4-5-20251101 (FC)Anthropic · Proprietary | 77.47% |
| 2 | Claude-Sonnet-4-5-20250929 (FC)Anthropic · Proprietary | 73.24% |
| 3 | Gemini-3-Pro-Preview (Prompt)Google · Proprietary | 72.51% |
| 4 | GLM-4.6 (FC thinking)Zhipu AI · MIT | 72.38% |
| 5 | Grok-4-1-fast-reasoning (FC)xAI · Proprietary | 69.57% |
Official score source · Retained source extract · Explore this track
EvalPlus · HumanEval+
Greedy decoding; HumanEval+ 0.1.10 and MBPP+ 0.2.0 (399 hand-verified MBPP tasks). Chat and completion modes are explicitly noted.
All source rows with this metric
| Position | Model / system | Pass@1 (%) |
|---|---|---|
| 1 | O1 Mini (Sept 2024)Chat prompting | 89% |
| 1 | O1 Preview (Sept 2024)Chat prompting | 89% |
| 3 | GPT 4o (Aug 2024)Chat prompting | 87.2% |
| 3 | Qwen2.5-Coder-32B-InstructChat prompting | 87.2% |
| 5 | DeepSeek-V3 (Nov 2024)Chat prompting | 86.6% |
Official score source · Retained source extract · Explore this track
EvalPlus · MBPP+
Greedy decoding; HumanEval+ 0.1.10 and MBPP+ 0.2.0 (399 hand-verified MBPP tasks). Chat and completion modes are explicitly noted.
All source rows with this metric
| Position | Model / system | Pass@1 (%) |
|---|---|---|
| 1 | O1 Preview (Sept 2024)Chat prompting | 80.2% |
| 2 | O1 Mini (Sept 2024)Chat prompting | 78.8% |
| 3 | Qwen2.5-Coder-32B-InstructChat prompting | 77% |
| 4 | DeepSeek-Coder-V2-InstructChat prompting | 75.1% |
| 5 | Gemini 1.5 Pro 002Chat prompting | 74.6% |
Official score source · Retained source extract · Explore this track
MMLU-Pro · overall
Official TIGER-Lab submission table. Source fractions are converted to percentages; source-provided evaluator or Self-Reported labels are retained.
All source rows with this metric
| Position | Model / system | Accuracy (%) |
|---|---|---|
| 1 | Gemini-3.1-ProTIGER-Lab | 91.16% |
| 2 | Gemini-3-Pro(11/25)Self-Reported | 90.1% |
| 3 | GPT-o1Self-Reported | 89.3% |
| 4 | Claude-4.6-Opus(Thinking)Self-Reported | 89.1% |
| 5 | Gemini-3-Flash(12/25)Self-Reported | 88.6% |
Official score source · Retained source extract · Explore this track
MMMU-Pro · overall
Official MMMU-Pro · overall column; human expert baselines excluded. Original, Pro, Vision and test scores are not merged.
All non-human source rows with this metric
| Position | Model / system | Overall accuracy (%) |
|---|---|---|
| 1 | Chance Vision 1.5proprietary · author | 86.9% |
| 2 | GPT-5.4 Thinking w/ toolsproprietary · author | 82.1% |
| 3 | GPT-5.4 Thinking w/o toolsproprietary · author | 81.2% |
| 4 | Gemini 3.0 Proproprietary · author | 81% |
| 5 | Gemini 3.1 Pro Thinking (High)proprietary · author | 80.5% |
Official score source · Retained source extract · Explore this track
MMMU · validation
Official MMMU · validation column; human expert baselines excluded. Original, Pro, Vision and test scores are not merged.
All non-human source rows with this metric
| Position | Model / system | Validation accuracy (%) |
|---|---|---|
| 1 | GPT-5.1proprietary · author | 85.4% |
| 2 | DreamPRM-1.5 (GPT-5-mini w/ thinking)proprietary · author | 84.6% |
| 3 | GPT-5 w/ thinkingproprietary · author | 84.2% |
| 4 | Gemini 2.5 Pro Deep-Thinkproprietary · author | 84% |
| 5 | o3proprietary · author | 82.9% |
Official score source · Retained source extract · Explore this track
Humanity’s Last Exam · accuracy
Official homepage model table; underlying prompting and tool settings must be checked at the source. Calibration error is displayed separately.
All source rows with this metric
| Position | Model / system | Accuracy (%) |
|---|---|---|
| 1 | Gemini 3 ProCalibration error 57.2% (lower is better) | 38.3% |
| 2 | GPT-5Calibration error 50.0% (lower is better) | 25.3% |
| 3 | Grok 4Calibration error 56.4% (lower is better) | 24.5% |
| 4 | Gemini 2.5 ProCalibration error 72.0% (lower is better) | 21.6% |
| 5 | GPT-5-miniCalibration error 65.0% (lower is better) | 19.4% |
Official score source · Retained source extract · Explore this track
τ³-bench · banking text
Banking text with knowledge retrieval. Official homepage top-three preview only; this is not the full leaderboard.
Official homepage top three only
| Position | Model / system | Pass^1 (%) |
|---|---|---|
| 1 | Qwen 3.8 MaxQwen | 55.2% |
| 2 | Claude Opus 5Anthropic | 48.7% |
| 3 | Grok 4.5xAI | 47.9% |
Official score source · Retained source extract · Explore this track
τ³-bench · voice
Voice across retail, airline, telecom and banking. Official homepage top-three preview only; this is not the full leaderboard.
Official homepage top three only
| Position | Model / system | Pass^1 (%) |
|---|---|---|
| 1 | gpt-live-1OpenAI | 81.7% |
| 2 | Pine Voice PreviewPine AI | 75.4% |
| 3 | grok-voice-think-fast-1.0xAI | 67.3% |
Official score source · Retained source extract · Explore this track
τ²-bench · text
Legacy τ² text across retail, airline and telecom; separate from the updated τ³ tasks. Official homepage top-three preview only; this is not the full leaderboard.
Official homepage top three only
| Position | Model / system | Pass^1 (%) |
|---|---|---|
| 1 | Qwen3.5-397B-A17BAlibaba Cloud | 87.9% |
| 2 | Gemini 3.0 ProGoogle | 85.4% |
| 3 | Claude Opus 4.5Anthropic | 85.3% |
Official score source · Retained source extract · Explore this track
LiveCodeBench · 2024-08-01 to 2025-05-01
Default official problem window 2024-08-01 through 2025-05-01. Scores averaged and rounded as in upstream; its eligibility-date filter is retained. Source dates include placeholders and are not independently verified model release dates or contamination evidence. This dated window is not a current frontier ranking.
Source models passing the upstream eligibility-date filter with scored problems in this window
| Position | Model / system | Pass@1 (%) |
|---|---|---|
| 1 | O4-Mini (High)454 problems · source eligibility date 2023-04-30 (unverified release date) | 80.2% |
| 2 | O3 (High)454 problems · source eligibility date 2023-04-30 (unverified release date) | 75.8% |
| 3 | O4-Mini (Medium)454 problems · source eligibility date 2023-04-30 (unverified release date) | 74.2% |
| 4 | Gemini-2.5-Pro-06-05454 problems · source eligibility date 2023-04-30 (unverified release date) | 73.6% |
| 5 | DeepSeek-R1-0528454 problems · source eligibility date 2024-06-30 (unverified release date) | 73.1% |
Official score source · Retained source extract · Explore this track
MathArena · expected performance
Organizer’s aggregate across non-deprecated competitions, with questions equally weighted. Source uncertainty and expected costs are retained. This is a four-row recommendation excerpt, not all models or a BuilderWars composite.
Official homepage recommendations only
| Position | Model / system | Expected normalized performance (%) |
|---|---|---|
| 1 | GPT-6.1 Sol (max)#1 overall · OpenAI · source uncertainty ±2.3 percentage points · expected cost $0.55 ±$0.045 | 90.1% |
| 2 | GPT-6 Astra (max)#2 overall · OpenAI · source uncertainty ±2.0 percentage points · expected cost $2.42 ±$0.28 | 88% |
| 3 | GPT-6 Sol (max)#3 overall · OpenAI · source uncertainty ±2.5 percentage points · expected cost $1.57 ±$0.17 | 84.9% |
| — | Qwen3.8-MaxBest open model · Qwen · source uncertainty ±7.6 percentage points · expected cost $3.07 ±$0.35 | 56.1% |
Official score source · Retained source extract · Explore this track
Tracks are not combined into a universal model ranking.