The paper’s authors propose an on-policy expert-correction pipeline, a scheme that finds the failing turn in the weaker model’s own rollout and asks the expert to rewrite only that turn.
Across seven enterprise agent tasks, training a weaker model on full expert trajectories after agent-harness evolution reduced performance by 4 to 30 points for Qwen3-Coder and Gemma 4. An agent harness is the system prompt, tool set, execution hooks, and context-management scaffolding around a model.
Claim check:
- The authors propose an on-policy expert-correction pipeline that finds the failing turn in the weaker model’s own rollout and asks the expert to rewrite only that turn. (confirmed by the publication itself: evidence; «We therefore develop an on-policy expert-correction pipeline, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model’s own rollout and asks the expert to rewrite only that turn.»)
- Across seven enterprise agent tasks, training a weaker model on full expert trajectories after agent-harness evolution reduced performance by 4 to 30 points for Qwen3-Coder and Gemma 4. (confirmed by the publication itself: evidence; «However, training the weaker model on the expert’s complete trajectories under the evolved harness backfires: performance regresses on all seven tasks by 4 to 30 points across Qwen3-Coder and Gemma 4, even though the same procedure helps under the unevolved harness.»)
- An agent harness includes the system prompt, tool set, execution hooks, and context-management scaffolding around a model. (confirmed by the publication itself: evidence; «Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success.»)
Primary sources:
score 62.4 · kind research · revision 1 · stories st-727nfl