In the preprint When Models Edit Too Much, Tongyao Zhu, Wei Hern Lim, and Min-Yen Kan examine how language models rewrite more code than needed to fix a bug. The work shows that repair quality includes minimality, reviewability, and fidelity to the original implementation.
The authors collected 400 BigCodeBench tasks and introduced controlled AST corruptions into the reference solutions. This produced a known minimal patch for each task.
Even for GPT-5.5, high Pass@1 came with overly large edits and added cognitive complexity. An instruction to preserve the original implementation reduced the mean excess Levenshtein distance from 0.195 to 0.131.
The same instruction reduced added cognitive complexity by 26.6% and increased Pass@1 by 2.3 points. The authors note that these improvements are not explained only by a larger reasoning budget or larger models.
Claim check:
- arXiv hosts a research preprint by Tongyao Zhu, Wei Hern Lim, and Min-Yen Kan on language models’ tendency to rewrite more code than needed to fix a bug. (confirmed by the publication itself: evidence; «Authors: Tongyao Zhu , Wei Hern Lim , Min-Yen Kan View a PDF of the paper titled When Models Edit Too Much: On the Fidelity of Minimal Code Edits, by Tongyao Zhu and 2 other authors View PDF HTML (experimental) Abstract: Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough: useful repairs should also be minimal, reviewable, and faithful to the original implementation. We study over-editing, the tendency of a model to rewrite code beyond what is required to fix a bug.»)
- The authors argue that a useful model-generated code repair should be not only correct but also minimal, reviewable, and faithful to the original implementation. (confirmed by the publication itself: evidence; «Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough: useful repairs should also be minimal, reviewable, and faithful to the original implementation.»)
- To assess edit minimality, the authors collected 400 BigCodeBench tasks, introduced controlled AST corruptions into the reference solutions, and thus obtained a known minimal patch for each task. (confirmed by the publication itself: evidence; «We construct an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch.»)
- Excessive editing is common in frontier LLMs: even for GPT-5.5, high Pass@1 can come with unjustifiably large edits and added cognitive complexity. (confirmed by the publication itself: evidence; «Across frontier LLMs, over-editing is widespread even among strong models like GPT-5.5: high Pass@1 can coexist with unnecessarily large edits and added cognitive complexity.»)
- An instruction to preserve the original implementation reduced the mean excess Levenshtein distance from 0.195 to 0.131, reduced added cognitive complexity by 26.6%, and increased Pass@1 by 2.3 points at the same time. (confirmed by the publication itself: evidence; «A preservation instruction substantially reduces this behavior, lowering average excess Levenshtein distance from 0.195 to 0.131, reducing added cognitive complexity by 26.6%, and increasing Pass@1 by 2.3 points.»)
- These improvements are not explained simply by a larger reasoning budget or larger models. (confirmed by the publication itself: evidence; «However, these gains do not simply follow from a larger reasoning budget or larger models.»)
- During post-training, supervised fine-tuning overfit to familiar corruption types, while reinforcement learning produced a better trade-off between out-of-distribution edit fidelity and retaining performance. (confirmed by the publication itself: evidence; «We observe that supervised fine-tuning overfits to seen corruption patterns, whereas reinforcement learning gives the best out-of-domain edit-fidelity and performance-retention trade-off.»)
- The work treats edit fidelity as a separate dimension of code repair quality that can be measured and trained. (confirmed by the publication itself: evidence; «These results position edit fidelity as a distinct axis of code-repair quality and show that it can be measured and learned.»)
Primary sources:
score 66.5 · kind research · revision 1 · stories st-1wd2cwd