- axis
- agentic
- verification
- det-gold
- whatItTests
- Contamination-free binary reverse-engineering: 19 private, real-world-scale programs averaging 16.9K lines of code, combined with 44 in-house anti-analysis primitives to yield 262 binary instances and 1,572 deterministically graded tasks that an agent must solve by analyzing compiled binaries.
- saturation
- open
- patternLabFit
- A hard, deterministic-graded frontier with real headroom (best model 61.4% per instance, only 31.5% of instances fully solved), plus a clean capability-transfer finding (strong source-code security skill does not transfer to binary analysis) that cautions against assuming a pattern validated on one coding surface generalizes to adjacent ones. Single-agent evaluation only; no collaboration discriminator asserted.
- notes
- Submitted 2026-08-11, within the discovery window. Benchmark name confirmed verbatim from the abstract: 'we introduce SRE-Bench, the first realistic, contamination-free RE benchmark.' Evaluated on five frontier LLMs (GPT-5.6-sol, Claude-Opus-5, GPT-5.5, Grok-4.5, GLM-5.2). Distinct from every tracked entry: no tracked benchmark covers security/reverse-engineering of compiled binaries; the private-programs-plus-mutated-anti-analysis-primitives design is a contamination defense analogous to SWE-bench Pro's held-out repos but on a different surface.