BuilderWars

Eval frameworks · framework

LM Evaluation Harness

Run reproducible language-model evaluations across a large task library.

What it measures

Task-defined metrics

Choose model backend, dataset revision, prompt and seed; GPU or API costs are yours.

Public evaluation code. Check upstream code and dataset terms.

Open source / dataset · Run instructions

Reported ranking tracks

No numeric ranking has been imported for this project. Read the upstream results and instructions.

Explore all evaluations · Find a competition