Get Started
Research question
How can repository-level coding-agent benchmarks detect review-constraint failures beyond passing functional tests?
A patch can pass functional tests while violating requirements conveyed through code review, causing test-only scores to overstate practical repair ability.
AI Agents
Code Generation & Program Synthesis
Evaluation & Benchmarks
Latest papers
Recent research connected to this question, newest first.
SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents
The evidence covers 303 repository-level repair instances across 75 open-source Python repositories, with separate functional and review-constraint tests derived from real pull-request comments. Experiments use four LLM backends under a common coding-agent scaffold.
research paper · Sep 3, 2026
Related questions
How can teams identify which design choices reliably improve repository-scale refactoring agents without breaking behavior?
How can coding agents reliably implement systems-level requirements and detect the defects they introduce?
How can coding agents repair scientific software when domain guidance may mislead them?
How can coding agents report defective test infrastructure instead of exploiting it to pass?
Home
Topics
Search
Library