BuilderWars

BUILDERWARS EVALS

The eval universe. One place to start.

50 benchmarks and frameworks. Find the right test, read the reported rankings, then build a better contender.

Explore reported rankings · Find a fun competition

Tool use · benchmark

ToolSandbox

Test stateful tool use with intermediate milestones and conversational dependencies.

Milestone-based task success

Open source / dataset

Reasoning · benchmark

MMLU

Measure knowledge across academic and professional subjects.

Subject and aggregate accuracy

Open source / dataset

Reasoning · benchmark

GPQA

Test graduate-level science reasoning on expert-authored questions.

Accuracy (%)

Open source / dataset

Mathematics · benchmark

MATH

Evaluate competition-level mathematical problem solving.

Exact-answer accuracy

Open source / dataset

Long context · benchmark

RULER

Measure effective context length with controlled synthetic tasks.

Accuracy by context length

Open source / dataset

Instruction following · benchmark

IFEval

Check instruction-following constraints with deterministic verifiers.

Strict / loose instruction accuracy

Open source / dataset

Eval frameworks · framework

Inspect AI

Build agent, tool-use and model evaluations with reusable scoring and sandboxes.

Evaluation-defined scores

Open source / dataset

Eval frameworks · framework

LightEval

Evaluate models using configurable tasks and multiple inference backends.

Task-defined metrics

Open source / dataset

Eval frameworks · framework

OpenCompass

Manage broad language-model evaluation with reusable model and dataset configurations.

Dataset-specific scores

Open source / dataset

Eval frameworks · framework

Harbor

Run isolated agent evaluations and retain trial artifacts across benchmark environments.

Task-defined metrics

Open source / dataset

RAG & applications · framework

Ragas

Evaluate retrieval, faithfulness and answer quality in RAG applications.

Retrieval and generation metrics

Open source / dataset

RAG & applications · framework

Promptfoo

Compare prompts, applications and agents with assertions and regression tests.

Assertion pass rate / configured scores

Open source / dataset

RAG & applications · framework

DeepEval

Test LLM applications with configurable metrics and evaluation datasets.

Metric-specific scores

Open source / dataset

Safety · framework

Garak

Probe model weaknesses with reproducible scanners and detectors.

Probe / detector outcomes

Open source / dataset

Games · framework

OpenSpiel

Study competitive and cooperative agents across game environments and algorithms.

Game-specific outcomes

Open source / dataset

Directory reviewed 2026-10-07. Dataset access and licenses vary by project.