- axis
- coding
- verification
- det-test
- whatItTests
- 89 hand-crafted, human-verified command-line tasks spanning scientific computing, software engineering, ML, security, sysadmin and data science, each requiring an agent to complete an end-to-end terminal workflow (compiling, training, configuring, debugging) rather than emit a single patch, graded with task-specific tests.
- discriminates
[build-verify-reflect]
- saturation
- open
- patternLabFit
- Tasks require multi-step terminal workflows rather than a single patch, giving a longer-horizon build-verify-reflect loop than SWE-bench-style tasks. Straddles the coding and agentic axes (structurally closer to OSWorld/tau-bench); scoped here as a coding-agent benchmark since its tasks are software-engineering-workflow-centric, but flagged for the assembler's judgment on axis placement.
- notes
- Published January 2026, within the 'significant recent benchmark' window. Distinct from the SWE-bench family (single-repo patch against a test harness): tasks span full terminal workflows across arbitrary tools, not just a git diff against one repository. Axis straddle (coding vs agentic, cf. osworld) is called out explicitly for the assembler. Version churn (re-verified 2026-08-22): the official tbench.ai news feed lists Terminal-Bench 2.0 (2025-11-07), 'a harder, better verified version of Terminal-Bench and a new package evaluating and optimizing agents', and Terminal-Bench 2.1 (2026-05-06), 'a revision of Terminal-Bench 2.0 that fixes 28 tasks and introduces continuous validation for agentic benchmarks'; the same feed also lists further variants (a Terminal-Bench 3.0 call for tasks Mar 2026, Terminal-Bench-Science May 2026, Terminal-Bench Challenges Jun 2026, Harbor-Index Jul 2026), so the tracked 89-task citation describes the original registry, not the currently versioned task set.