- axis
- agentic
- verification
- det-gold
- whatItTests
- 466 real-world questions requiring reasoning, multimodal understanding, and tool use, each with a short, exact-match gold answer; 300 answers held back to power a private leaderboard.
- saturation
- open
- patternLabFit
- Exact-match gold answers make it usable as an outcome benchmark for any pattern, but the paper itself only evaluates single-agent assistants (GPT-4 + plugins), so no specific collaboration-pattern fit is evidenced.
- notes
- Not new (2023) but widely cited and missing; large human/AI gap at publication.