- axis
- coding
- verification
- det-test
- whatItTests
- Continuously refreshed competitive-programming problems scraped from LeetCode, AtCoder and CodeForces after each evaluated model's training cutoff, testing code generation plus self-repair, code execution, and test-output prediction.
- discriminates
[build-verify-reflect]
- saturation
- mid
- patternLabFit
- The self-repair task type is a native build-verify-reflect probe (write, run, read failure, revise) with an execution oracle. The paper evaluates single models only, so no multi-agent collaboration pattern is evidenced directly by the source.
- notes
- Distinct from swe-bench-verified/swe-bench-pro (real GitHub issues) and the CodeContests/USACO entries proposed alongside it (fixed problem sets): LiveCodeBench's defining feature is a rolling window of post-cutoff competitive-programming problems plus explicit self-repair/execution/test-output-prediction task types, not just generation accuracy. Saturation marked mid rather than a precise figure: the paper itself gives no aggregate ceiling, and current third-party leaderboard chatter about near-saturation at the frontier is not an authoritative source, so no specific score is asserted here.