Reasoning · benchmark
BIG-Bench Hard
Probe reasoning through challenging tasks selected from BIG-Bench.
What it measures
Task accuracy
Use a declared prompting protocol; this is not all of BIG-Bench.
Public evaluation code. Check upstream code and dataset terms.
Open source / dataset · Run instructions
Reported ranking tracks
No numeric ranking has been imported for this project. Read the upstream results and instructions.