Reasoning · benchmark
MMLU-Pro
Evaluate broad knowledge and reasoning with more challenging multiple-choice questions.
What it measures
Overall accuracy (%)
Chain-of-thought and answer-extraction settings affect comparability.
Public evaluation code. Check upstream code and dataset terms.
Open source / dataset · Run instructions · Official results
Reported ranking tracks
- MMLU-Pro · overall · 262 entries · captured 2026-10-07T07:01:28Z