BuilderWars

Long context · benchmark

LongBench

Evaluate long-context understanding across tasks and languages.

What it measures

Task-specific score / V2 accuracy

Original LongBench and LongBench V2 use different tasks and metrics.

Public evaluation code. Check upstream code and dataset terms.

Open source / dataset · Run instructions · Official results

Reported ranking tracks

No numeric ranking has been imported for this project. Read the upstream results and instructions.

Explore all evaluations · Find a competition