BuilderWars

REPORTED RANKINGS

Know the score. Know the conditions.

14 ranking tracks. Captured 2026-10-07T07:01:28Z. Scores are reported by the linked benchmark organizers. BuilderWars republishes dated source extracts and does not independently run or certify these evaluations.

SWE-bench Verified · agent systems

500 Verified tasks; Different submitted agent systems and model configurations; this is a system ranking.

All official Verified submissions with a reported score

SWE-bench Verified · agent systems · Resolved tasks (%)
PositionModel / systemResolved tasks (%)
1live-SWE-agent + Claude 4.5 Opus medium (20251101)live-SWE-agent · Submission · Reasoning medium · System: Attempts - 179.2%
1Sonar Foundation Agent + Claude 4.5 OpusSonar Foundation Agent · Submission · System: Attempts - 179.2%
3TRAE + Doubao-Seed-CodeTRAE · Submission · System: Attempts - 2+78.8%
4live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)live-SWE-agent · Submission · System: Attempts - 177.4%
5Atlassian Rovo Dev (2025-09-02)Atlassian Rovo Dev · Submission · System: Attempts - 276.8%

Official score source · Retained source extract · Explore this track

SWE-bench Verified · mini-SWE-agent

500 Verified tasks; mini-SWE-agent submissions only, including declared agent versions and settings.

mini-SWE-agent subset of official Verified data

SWE-bench Verified · mini-SWE-agent · Resolved tasks (%)
PositionModel / systemResolved tasks (%)
1Claude 4.5 Opus (high)mini-SWE-agent · Submission · Harness version 2.0.0 · Reasoning high · System: Attempts - 176.8%
2Gemini 3 Flash (high)mini-SWE-agent · Submission · Harness version 2.0.0 · Reasoning high · System: Attempts - 175.8%
2MiniMax M2.5 (high)mini-SWE-agent · Submission · Harness version 2.0.0 · Reasoning high · System: Attempts - 175.8%
4Claude 4.6 Opusmini-SWE-agent · Submission · Harness version 2.0.0 · System: Attempts - 175.6%
5Claude 4.5 Opus (20251101) (medium)mini-SWE-agent · Organizer-checked · Harness version 1.16.0 · Reasoning medium · System: Attempts - 174.4%

Official score source · Retained source extract · Explore this track

BFCL V4 · overall

V4 overall score; FC and Prompt calling modes remain in model names. Estimated benchmark cost/latency are not independent billing evidence.

All source rows with this metric

BFCL V4 · overall · Overall accuracy (%)
PositionModel / systemOverall accuracy (%)
1Claude-Opus-4-5-20251101 (FC)Anthropic · Proprietary77.47%
2Claude-Sonnet-4-5-20250929 (FC)Anthropic · Proprietary73.24%
3Gemini-3-Pro-Preview (Prompt)Google · Proprietary72.51%
4GLM-4.6 (FC thinking)Zhipu AI · MIT72.38%
5Grok-4-1-fast-reasoning (FC)xAI · Proprietary69.57%

Official score source · Retained source extract · Explore this track

EvalPlus · HumanEval+

Greedy decoding; HumanEval+ 0.1.10 and MBPP+ 0.2.0 (399 hand-verified MBPP tasks). Chat and completion modes are explicitly noted.

All source rows with this metric

EvalPlus · HumanEval+ · Pass@1 (%)
PositionModel / systemPass@1 (%)
1O1 Mini (Sept 2024)Chat prompting89%
1O1 Preview (Sept 2024)Chat prompting89%
3GPT 4o (Aug 2024)Chat prompting87.2%
3Qwen2.5-Coder-32B-InstructChat prompting87.2%
5DeepSeek-V3 (Nov 2024)Chat prompting86.6%

Official score source · Retained source extract · Explore this track

EvalPlus · MBPP+

Greedy decoding; HumanEval+ 0.1.10 and MBPP+ 0.2.0 (399 hand-verified MBPP tasks). Chat and completion modes are explicitly noted.

All source rows with this metric

EvalPlus · MBPP+ · Pass@1 (%)
PositionModel / systemPass@1 (%)
1O1 Preview (Sept 2024)Chat prompting80.2%
2O1 Mini (Sept 2024)Chat prompting78.8%
3Qwen2.5-Coder-32B-InstructChat prompting77%
4DeepSeek-Coder-V2-InstructChat prompting75.1%
5Gemini 1.5 Pro 002Chat prompting74.6%

Official score source · Retained source extract · Explore this track

MMLU-Pro · overall

Official TIGER-Lab submission table. Source fractions are converted to percentages; source-provided evaluator or Self-Reported labels are retained.

All source rows with this metric

MMLU-Pro · overall · Accuracy (%)
PositionModel / systemAccuracy (%)
1Gemini-3.1-ProTIGER-Lab91.16%
2Gemini-3-Pro(11/25)Self-Reported90.1%
3GPT-o1Self-Reported89.3%
4Claude-4.6-Opus(Thinking)Self-Reported89.1%
5Gemini-3-Flash(12/25)Self-Reported88.6%

Official score source · Retained source extract · Explore this track

MMMU-Pro · overall

Official MMMU-Pro · overall column; human expert baselines excluded. Original, Pro, Vision and test scores are not merged.

All non-human source rows with this metric

MMMU-Pro · overall · Overall accuracy (%)
PositionModel / systemOverall accuracy (%)
1Chance Vision 1.5proprietary · author86.9%
2GPT-5.4 Thinking w/ toolsproprietary · author82.1%
3GPT-5.4 Thinking w/o toolsproprietary · author81.2%
4Gemini 3.0 Proproprietary · author81%
5Gemini 3.1 Pro Thinking (High)proprietary · author80.5%

Official score source · Retained source extract · Explore this track

MMMU · validation

Official MMMU · validation column; human expert baselines excluded. Original, Pro, Vision and test scores are not merged.

All non-human source rows with this metric

MMMU · validation · Validation accuracy (%)
PositionModel / systemValidation accuracy (%)
1GPT-5.1proprietary · author85.4%
2DreamPRM-1.5 (GPT-5-mini w/ thinking)proprietary · author84.6%
3GPT-5 w/ thinkingproprietary · author84.2%
4Gemini 2.5 Pro Deep-Thinkproprietary · author84%
5o3proprietary · author82.9%

Official score source · Retained source extract · Explore this track

Humanity’s Last Exam · accuracy

Official homepage model table; underlying prompting and tool settings must be checked at the source. Calibration error is displayed separately.

All source rows with this metric

Humanity’s Last Exam · accuracy · Accuracy (%)
PositionModel / systemAccuracy (%)
1Gemini 3 ProCalibration error 57.2% (lower is better)38.3%
2GPT-5Calibration error 50.0% (lower is better)25.3%
3Grok 4Calibration error 56.4% (lower is better)24.5%
4Gemini 2.5 ProCalibration error 72.0% (lower is better)21.6%
5GPT-5-miniCalibration error 65.0% (lower is better)19.4%

Official score source · Retained source extract · Explore this track

τ³-bench · banking text

Banking text with knowledge retrieval. Official homepage top-three preview only; this is not the full leaderboard.

Official homepage top three only

τ³-bench · banking text · Pass^1 (%)
PositionModel / systemPass^1 (%)
1Qwen 3.8 MaxQwen55.2%
2Claude Opus 5Anthropic48.7%
3Grok 4.5xAI47.9%

Official score source · Retained source extract · Explore this track

τ³-bench · voice

Voice across retail, airline, telecom and banking. Official homepage top-three preview only; this is not the full leaderboard.

Official homepage top three only

τ³-bench · voice · Pass^1 (%)
PositionModel / systemPass^1 (%)
1gpt-live-1OpenAI81.7%
2Pine Voice PreviewPine AI75.4%
3grok-voice-think-fast-1.0xAI67.3%

Official score source · Retained source extract · Explore this track

τ²-bench · text

Legacy τ² text across retail, airline and telecom; separate from the updated τ³ tasks. Official homepage top-three preview only; this is not the full leaderboard.

Official homepage top three only

τ²-bench · text · Pass^1 (%)
PositionModel / systemPass^1 (%)
1Qwen3.5-397B-A17BAlibaba Cloud87.9%
2Gemini 3.0 ProGoogle85.4%
3Claude Opus 4.5Anthropic85.3%

Official score source · Retained source extract · Explore this track

LiveCodeBench · 2024-08-01 to 2025-05-01

Default official problem window 2024-08-01 through 2025-05-01. Scores averaged and rounded as in upstream; its eligibility-date filter is retained. Source dates include placeholders and are not independently verified model release dates or contamination evidence. This dated window is not a current frontier ranking.

Source models passing the upstream eligibility-date filter with scored problems in this window

LiveCodeBench · 2024-08-01 to 2025-05-01 · Pass@1 (%)
PositionModel / systemPass@1 (%)
1O4-Mini (High)454 problems · source eligibility date 2023-04-30 (unverified release date)80.2%
2O3 (High)454 problems · source eligibility date 2023-04-30 (unverified release date)75.8%
3O4-Mini (Medium)454 problems · source eligibility date 2023-04-30 (unverified release date)74.2%
4Gemini-2.5-Pro-06-05454 problems · source eligibility date 2023-04-30 (unverified release date)73.6%
5DeepSeek-R1-0528454 problems · source eligibility date 2024-06-30 (unverified release date)73.1%

Official score source · Retained source extract · Explore this track

MathArena · expected performance

Organizer’s aggregate across non-deprecated competitions, with questions equally weighted. Source uncertainty and expected costs are retained. This is a four-row recommendation excerpt, not all models or a BuilderWars composite.

Official homepage recommendations only

MathArena · expected performance · Expected normalized performance (%)
PositionModel / systemExpected normalized performance (%)
1GPT-6.1 Sol (max)#1 overall · OpenAI · source uncertainty ±2.3 percentage points · expected cost $0.55 ±$0.04590.1%
2GPT-6 Astra (max)#2 overall · OpenAI · source uncertainty ±2.0 percentage points · expected cost $2.42 ±$0.2888%
3GPT-6 Sol (max)#3 overall · OpenAI · source uncertainty ±2.5 percentage points · expected cost $1.57 ±$0.1784.9%
—Qwen3.8-MaxBest open model · Qwen · source uncertainty ±7.6 percentage points · expected cost $3.07 ±$0.3556.1%

Official score source · Retained source extract · Explore this track

Tracks are not combined into a universal model ranking.