The authors introduced ExecCritic, a method for agents that write code. It separates test creation from source-code repair.

One agent creates tests, while a validation mechanism that stops on errors accepts and freezes them. Another agent repairs code from test execution feedback and does not change the tests.

On SWE-bench Verified, the combined post-trained agents reached 72.6% resolved tasks. The authors report this is 11.4 percentage points above the original no-test baseline.

Claim check:

  • The authors introduced ExecCritic, a method for agents that write code. (confirmed by the publication itself: evidence; «We introduce ExecCritic, combining a test–verify–revise scaffold with a role-specific reinforcement learning recipe for training agents within it.»)
  • It separates test creation from source-code repair. (confirmed by the publication itself: evidence; «The scaffold separates test construction from source-code repair»)
  • One agent creates tests, while a validation mechanism that stops on errors accepts and freezes them. (confirmed by the publication itself: evidence; «a Test agent independently generates repository-native tests, a fail-closed harness qualifies and freezes them»)
  • Another agent repairs code from test execution feedback and does not change the tests. (confirmed by the publication itself: evidence; «a Repair agent revises source code from their execution feedback without changing the tests.»)
  • On SWE-bench Verified, the combined post-trained agents reached 72.6% resolved tasks. (confirmed by the publication itself: evidence; «composing the two post-trained Qwen agents reaches 72.6%»)
  • The authors report this is 11.4 percentage points above the original no-test baseline. (confirmed by the publication itself: evidence; «an 11.4-point gain over the original no-test baseline»)

Primary sources:

score 72.6 · kind research · revision 1 · stories st-zfcxb4