Eval frameworks · framework
LM Evaluation Harness
Run reproducible language-model evaluations across a large task library.
What it measures
Task-defined metrics
Choose model backend, dataset revision, prompt and seed; GPU or API costs are yours.
Public evaluation code. Check upstream code and dataset terms.
Open source / dataset · Run instructions
Reported ranking tracks
No numeric ranking has been imported for this project. Read the upstream results and instructions.