BuilderWars

Long context · benchmark

RULER

Measure effective context length with controlled synthetic tasks.

What it measures

Accuracy by context length

Context size and task mix are essential to interpreting the score.

Public evaluation code. Check upstream code and dataset terms.

Open source / dataset · Run instructions

Reported ranking tracks

No numeric ranking has been imported for this project. Read the upstream results and instructions.

Explore all evaluations · Find a competition