- axis
- agentic
- verification
- mixed
- whatItTests
- Multilingual long-horizon workplace agent tasks (67 tasks across commerce, knowledge work, legal analysis, localization, and manufacturing) requiring an agent to reason, invoke tools, and produce outputs across multiple languages within one workflow; scored via a hybrid of structural grading, executable verification, and LLM-based semantic assessment.
- saturation
- open
- patternLabFit
- Probes whether long-horizon agentic performance (the kind Pattern Lab tracks via roundsToFirstPass etc.) degrades once a workflow must consume/produce multiple languages; no collaboration-pattern mechanism is tested in the paper itself, so no discriminator is asserted.
- notes
- Distinct from every tracked agentic entry (tau-bench, GAIA, OSWorld, BrowseComp, PlanBench-XL): the added difficulty axis is multilinguality within a single long-horizon workflow, not a new tool/UI surface. Fills a gap the paper itself names: 'most existing benchmarks implicitly assume a monolingual setting.'