Eval frameworks · framework
VLMEvalKit
Run multimodal benchmarks through a common evaluation toolkit.
What it measures
Benchmark-specific metrics
Dataset downloads, model access and judge settings vary.
Public evaluation code. Check upstream code and dataset terms.
Open source / dataset · Run instructions
Reported ranking tracks
No numeric ranking has been imported for this project. Read the upstream results and instructions.