Long context · benchmark
LongBench
Evaluate long-context understanding across tasks and languages.
What it measures
Task-specific score / V2 accuracy
Original LongBench and LongBench V2 use different tasks and metrics.
Public evaluation code. Check upstream code and dataset terms.
Open source / dataset · Run instructions · Official results
Reported ranking tracks
No numeric ranking has been imported for this project. Read the upstream results and instructions.