The authors introduced SWE-Gate, a repository benchmark for software engineering agents. It evaluates compliance with review constraints alongside functional correctness.

SWE-Gate extracts review constraints from comments on real pull requests. The benchmark then creates repository-level code-fixing tasks around them.

Each task includes separate functional tests and constraint tests. Each task also includes non-compliant and reference patches, so it can separately assess problem fixing and compliance with review requirements.

In the experiments, 4 LLM backends ran in a shared coding-agent framework. Of 644 fixes that passed the functional tests, 221 did not meet the review constraints.

Claim check:

  • The paper’s authors introduced SWE-Gate, a repository benchmark for software engineering agents that evaluates compliance with review constraints alongside functional correctness. (confirmed by the publication itself: evidence; «We introduce SWE-Gate, a repository-level benchmark for software engineering agents that explicitly evaluates review constraint compliance alongside functional correctness.»)
  • SWE-Gate extracts review constraints from comments on real pull requests and creates repository-level code-fixing tasks around them. (confirmed by the publication itself: evidence; «SWE-Gate derives review constraints from real pull request review comments and synthesizes repository-level repair instances around these constraints.»)
  • Each benchmark task includes separate functional and constraint tests, as well as non-compliant and reference patches, allowing separate assessment of problem fixing and compliance with review requirements. (confirmed by the publication itself: evidence; «Each instance provides separate functional and constraint tests, together with non-compliant and gold patches, enabling explicit separation between issue resolution capability and review constraint compliance.»)
  • The SWE-Gate dataset includes 303 repository-level code-fixing tasks from 75 open Python repositories across different domains. (confirmed by the publication itself: evidence; «We construct SWE-Gate with 303 repository-level repair instances spanning 75 open-source Python repositories across diverse software domains.»)
  • The experiments used four LLM backends with different capability levels in a shared coding-agent framework. (confirmed by the publication itself: evidence; «Experiments with four LLM backends spanning different capability levels under a common coding-agent scaffold reveal a substantial gap between functional success and success under the complete repair specification: among 644 repairs that pass the functional tests, 221 fail to satisfy the provided review constraints.»)
  • Of 644 fixes that passed the functional tests, 221 did not meet the specified review constraints. (confirmed by the publication itself: evidence; «among 644 repairs that pass the functional tests, 221 fail to satisfy the provided review constraints.»)
  • The authors conclude that evaluation based only on functional tests overestimates agents’ ability to meet the full set of requirements for repository-level code-fixing tasks. (confirmed by the publication itself: evidence; «These findings show that functional-only evaluation overestimates agents’ ability to satisfy the full requirements of repository-level repair tasks.»)

Primary sources:

score 64.3 · kind research · revision 1 · stories st-1vvofgw