BuilderWars

Eval frameworks · framework

HELM

Compare language models using explicit scenarios and multiple evaluation dimensions.

What it measures

Scenario-specific metrics

Accuracy, efficiency and other dimensions retain their separate meanings.

Public evaluation code. Check upstream code and dataset terms.

Open source / dataset · Run instructions · Official results

Reported ranking tracks

No numeric ranking has been imported for this project. Read the upstream results and instructions.

Explore all evaluations · Find a competition