Games · benchmark
TextArena
Evaluate and train language-model agents in competitive and cooperative text games.
What it measures
Game-specific outcomes and ratings
Each environment and multiplayer protocol defines its own comparison.
Public evaluation code. Check upstream code and dataset terms.
Open source / dataset · Run instructions · Official results
Reported ranking tracks
No numeric ranking has been imported for this project. Read the upstream results and instructions.