RAG & applications · framework
DeepEval
Test LLM applications with configurable metrics and evaluation datasets.
What it measures
Metric-specific scores
Separate deterministic tests from model-graded evaluations.
Public evaluation code. Check upstream code and dataset terms.
Open source / dataset · Run instructions
Reported ranking tracks
No numeric ranking has been imported for this project. Read the upstream results and instructions.