- axis
- multi-agent
- verification
- mixed
- whatItTests
- Evaluates LLMs in pure-coordination (no competition) settings via two tasks: Agentic Coordination, where LLMs act as participants in four coordination games, and CoordQA, 198 multiple-choice questions testing Environment Comprehension, Theory-of-Mind reasoning, and Joint Planning; includes zero-shot coordination with unseen partners.
- saturation
- mid
- patternLabFit
- The CoordQA multiple-choice split (Environment Comprehension / ToM Reasoning / Joint Planning) gives a deterministic, decomposed diagnostic for exactly the sub-skills Pattern Lab's coordination-topology patterns rely on, useful for isolating why a topology underperforms rather than only measuring end-task success.
- notes
- Distinct from MultiAgentBench's topology comparison and BattleAgentBench's staged cooperation/competition ladder: LLM-Coordination targets pure coordination only (no adversarial component) and decouples in-game performance from a separate multiple-choice ToM/joint-planning diagnostic, plus zero-shot coordination with unseen partners. No discriminates tag confidently maps to the valid list from the abstract alone, left empty per no-fabrication policy.