- axis
- agentic
- verification
- det-gold
- whatItTests
- 1,266 hard-to-find, entangled trivia-style questions requiring persistent multi-step web browsing to locate a single short, exactly-verifiable answer.
- saturation
- mid
- patternLabFit
- Short verifiable answers make roundsToFirstPass and pass@k straightforward to compute for browsing agents; the paper compares single models/agents (GPT-4o, o1, Deep Research) and does not test collaboration patterns.
- notes
- OpenAI-authored; positioned as analogous to competitive programming for browsing agents. Saturation update (2026-07-17): OpenAI's GPT-5.6 Sol (released 2026-07-09, after the 2026-07-01 baseline) posts a new BrowseComp state-of-the-art of 92.2%. Independent leaderboard data corroborates a tight cluster of frontier models in the high-80s to low-90s on the same benchmark: standalone Claude Fable 5 near 90.8%, a cost-optimized Fable 5 orchestrator with Sonnet 5 workers at 86.8%, and GPT-5.5 Pro at 90.1%. BrowseComp's original paper (2025) showed frontier models far below this (o1 ~9.9%, best 'Deep Research' agent 51.5%). Marked 'mid' rather than 'saturated': this is OpenAI's own self-reported figure with no independent third-party leaderboard re-run confirming the exact 92.2%, and the cluster still falls short of the near-total ceiling seen on already-'saturated' entries (MATH-500, GPQA-Diamond). Re-verification (2026-08-22): the 92.2% figure now also appears on the third-party aggregator BenchLM.ai (updated 2026-08-22, listing GPT-5.6 Sol 92.2%, Kimi K3 91.2%, and Claude Opus 5 90.8% and noting the top models are clustered within 1.4 points, 'suggesting this benchmark is nearing saturation for frontier models'); BenchLM aggregates publicly reported figures rather than running independent re-runs, so the no-independent-re-run caveat still stands and saturation stays 'mid'.