1,800 code-completion evaluation instances across six programming languages and six task categories, derived from real developer telemetry and synthesized using generator models from multiple provider families, scored with functional correctness, similarity metrics, and LLM-judge assessments of usefulness and contextual relevance.
discriminates
[build-verify-reflect][judge-calibration]
saturation
open
patternLabFit
Combines an execution/functional-correctness oracle (build-verify-reflect) with an LLM-judge layer scoring usefulness and contextual relevance, making it usable both as a hard-oracle benchmark and as a judge-calibration testbed for whatever judge scores the qualitative dimensions. The paper evaluates 9 individual models, not multi-agent collaboration.
Published January 2026, within the 'significant recent benchmark' window. Distinct from bigcodebench (synthetic library/function-call tasks) and evocodebench/repobench (repo-derived generation/completion): DevBench is grounded in real developer telemetry across six languages and mixes functional-correctness, similarity, and LLM-judge scoring in one framework. Not to be confused with the differently-named 'DevEval' benchmark (arXiv:2403.08604), a separate full-software-development-lifecycle case study; that paper was reviewed and excluded here to avoid conflating the two distinct works.