- axis
- coding
- verification
- det-test
- whatItTests
- 1,632 expert-annotated issue-resolving tasks spanning seven languages (Java, TypeScript, JavaScript, Go, Rust, C, C++), extending the SWE-bench issue-to-patch format beyond Python.
- discriminates
[build-verify-reflect]
- saturation
- open
- patternLabFit
- Applies the same build-verify-reflect (patch, run tests, revise) loop as SWE-bench Verified/Pro across seven languages, letting collaboration patterns tuned on Python be tested for transfer to other ecosystems. The paper evaluates single-agent frameworks (Agentless, SWE-agent, OpenHands) individually, not multi-agent patterns.
- notes
- Distinct from swe-bench-verified/swe-bench-pro (Python-only): tests the identical issue-to-patch mechanism across six additional language ecosystems, exposing whether a fix strategy generalizes beyond Python idioms and tooling. No aggregate resolve-rate figure is asserted here beyond what the abstract states, since the paper's own summary does not give one.