- axis
- agentic
- verification
- det-test
- whatItTests
- 108 long-horizon computer-use workflows spanning real desktop and web applications across professional domains, each an end-to-end task taking human users a median of about 1.6 hours and requiring an average of 318 tool calls (versus about 30 in the original OSWorld), designed to surface failures where agents lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, or skip verification.
- discriminates
[build-verify-reflect]
- saturation
- open
- patternLabFit
- Same execution-based per-task oracle as OSWorld 1.0 but stretched over far longer horizons, giving real headroom (best model completes only 20.6% of tasks / 54.8% partial credit at a 500-step budget). Direct successor for any Pattern Lab experiment currently relying on the now-saturated original OSWorld for a computer-use ceiling.
- notes
- Official successor to OSWorld (id: osworld), built by the same XLANG Lab maintainers, arXiv:2606.29537 (submitted 2026-06-26), public repo/leaderboard at github.com/xlang-ai/OSWorld-V2. Distinct from tracked osworld: roughly 11x more tool calls per task on average, professional-domain workflows rather than short discrete tasks, explicit design goal of resisting the saturation the original benchmark hit. Verbatim from the abstract: 'Each task represents a realistic end-to-end workflow that takes human users a median of about 1.6 hours to complete and requires an average of 318 tool calls with Claude Opus 4.7 using maximum thinking, compared with about 30 in OSWorld 1.0.' Best model (Claude Opus 4.8, max thinking) completes only 20.6% of tasks at 54.8% partial score. Release date (2026-06-26) is about three weeks before this lens's 2026-07-17 discovery-window start, but the tracked dataset has not yet captured it, and the material fact it establishes (original osworld is saturated per its own maintainers) is what motivates the accompanying update below.