Growing Harness moves recurring control decisions from a language model’s context into executable code. The method starts with a scaffold that has no controller for solving tasks.
Execution traces tie each failure to a bounded code area, and an optimizer repairs several failures together. A held-out-task check rolls back repair sequences that harm earlier capabilities.
On BrowseComp-Plus and WebArena-Verified, Growing Harness achieved the best mean result in five of six benchmark-model combinations and trailed by 0.7 percentage points in the sixth.
On WebArena-Verified, Growing Harness achieved 44.7–45.3% success across models from 4B to 120B, while a 4B Tool-Calling agent reached 6.7%. Compared with a Tool-Calling agent, Growing Harness reduced language-model calls by 76.0–91.8% and deployed-agent inference cost by 74.4–98.6%.
Claim check:
- Growing Harness moves recurring control decisions from a language model’s context into executable code. (confirmed by the publication itself: evidence; «These results show that persistent program growth can move recurring control out of model context and into low-cost code, yielding reusable specialist agents that remain effective with smaller deployment models.»)
- The method starts with a scaffold that has no controller for solving tasks. (confirmed by the publication itself: evidence; «We introduce Growing Harness, a failure-guided training paradigm that learns the agent harness itself from a strategy-free scaffold that exposes fixed model and tool interfaces but encodes no task-solving controller.»)
- Execution traces tie each failure to a bounded code area, and an optimizer repairs several failures together. (confirmed by the publication itself: evidence; «Function-level execution traces localize each failure to a bounded code surface, an optimizer repairs a window of failures jointly, and a success-first held-out gate rolls back repair sequences that harm prior capability.»)
- A held-out-task check rolls back repair sequences that harm earlier capabilities. (confirmed by the publication itself: evidence; «Function-level execution traces localize each failure to a bounded code surface, an optimizer repairs a window of failures jointly, and a success-first held-out gate rolls back repair sequences that harm prior capability.»)
- On BrowseComp-Plus and WebArena-Verified, Growing Harness achieved the best mean result in five of six benchmark-model combinations and trailed by 0.7 percentage points in the sixth. (confirmed by the publication itself: evidence; «Across BrowseComp-Plus and WebArena-Verified with three deployment models from 4B to 120B parameters, Growing Harness achieves the highest mean success in five of six benchmark-model settings and trails the best mean by 0.7 pp. in the sixth.»)
- On WebArena-Verified, Growing Harness achieved 44.7–45.3% success across models from 4B to 120B, while a 4B Tool-Calling agent reached 6.7%. (confirmed by the publication itself: evidence; «On WebArena-Verified, its success remains 44.7-45.3% across model scales, whereas Tool-Calling falls to 6.7% with the 4B model.»)
- Compared with a Tool-Calling agent, Growing Harness reduced language-model calls by 76.0–91.8% and deployed-agent inference cost by 74.4–98.6%. (confirmed by the publication itself: evidence; «Relative to a Tool-Calling agent, it reduces LLM calls by 76.0-91.8% and deployed-agent inference cost by 74.4-98.6%.»)
Primary sources:
score 63.4 out of 100 · kind: research