Reasoning · benchmark
ARC-AGI-2
Solve novel grid transformations with strict separation between training and evaluation.
What it measures
Solved tasks and cost per task
Base models, reasoning systems and competition systems use different constraints.
Public evaluation code. Check upstream code and dataset terms.
Open source / dataset · Run instructions · Official results
Reported ranking tracks
No numeric ranking has been imported for this project. Read the upstream results and instructions.