- Unit
- One row = one benchmark or evaluation framework, with the axis it measures, how candidate answers are verified, what it tests, which collaboration patterns or quality metrics it discriminates, its saturation at the frontier, and a defining citation.
- Scope
- Two kinds of entry: outcome benchmarks to run collaboration patterns on (math, coding, agentic, knowledge), and multi-agent benchmarks plus the methodology that validates collaboration-quality metrics (judge calibration, Bradley-Terry and ELO ranking, inter-rater reliability, process evaluation).
- Verification
- Deterministic gold answer, deterministic test suite, LLM judge, human annotation, or mixed. Frameworks that are methodology rather than a scored benchmark are marked as not scored.
- Saturation
- Whether the frontier has run out of headroom. A saturated benchmark no longer separates strong systems, which is the field to read first when choosing what to measure against.
- Sources
- Each row carries a defining citation, the paper or the official benchmark repository. URLs travel with the dataset.
- License
- Free to use, distribute, and reproduce with attribution to Hermosa Labs LLC (CC BY 4.0).
- Cite
- Hermosa Labs LLC, “Agent Evals and Benchmarks,” 2026-08-22.