BuilderWars

Games · benchmark

TextArena

Evaluate and train language-model agents in competitive and cooperative text games.

What it measures

Game-specific outcomes and ratings

Each environment and multiplayer protocol defines its own comparison.

Public evaluation code. Check upstream code and dataset terms.

Open source / dataset · Run instructions · Official results

Reported ranking tracks

No numeric ranking has been imported for this project. Read the upstream results and instructions.

Explore all evaluations · Find a competition