- axis
- agentic
- verification
- mixed
- whatItTests
- 10,400 graph-analysis tasks with verifiable ground truth across three graph types and four task categories (graph retrieval, graph theory, graph machine learning, graph open-ended question answering), where an agent must access graph data and execute operations through 84 executable tools.
- saturation
- open
- patternLabFit
- Its headline finding that harness choice significantly affects performance on complex graph tasks makes it one of the few benchmarks that explicitly discriminates between agent harnesses rather than only between models, a useful control for any Pattern Lab claim that a collaboration-pattern gain is not just a harness effect. The paper evaluates single agents; no multi-agent mechanism is tested, so no discriminator is asserted.
- notes
- Submitted 2026-08-03, two days before the 2026-08-05 window start but not captured by the previous run, so treated as new to the catalogue. Distinct from every tracked agentic entry (tau-bench, GAIA, OSWorld/2.0, BrowseComp, PlanBench-XL, DevicesWorld, PolyWorkBench, MAG, Relay-Bench, HANDBOOK.md): the task surface is programmatic graph analysis over executable graph tools, not web/OS/dialogue/workplace tasks. Verification corrected to mixed at validation: the paper's primary metric, task success rate (SR), is scored by an LLM-as-a-judge comparing each agent's final response against the ground-truth answer for semantic equivalence across all four task categories, while a secondary metric, Tool Selection Accuracy (TSA), is scored deterministically via longest-common-subsequence comparison against the gold tool-call sequence.