The authors of FIRE study runtime policies: natural-language instructions and action denials that an agent harness applies in states preceding observed failures. The authors say these policies do not change model weights or the user prompt.
Across the full 87-task Terminal-Bench 2.1 suite, repeated success for GPT-5.6 Sol rose from 64.4% to 73.6%. In a randomized five-arm experiment, real policies achieved 61% success on eligible tasks, while no policy achieved 39%.
Claim check:
- The authors of FIRE study runtime policies: targeted natural-language instructions and action denials that an agent harness applies in states before observed failures. (confirmed by the publication itself: evidence; «We study runtime policies: targeted natural-language instructions and action denials applied by the agent harness at states that preceded observed failures»)
- The authors write that these policies do not change model weights or the user prompt. (confirmed by the publication itself: evidence; «without changing model weights or the user prompt»)
- Across the full 87-task Terminal-Bench 2.1 suite, repeated success for GPT-5.6 Sol rose from 64.4% to 73.6%. (confirmed by the publication itself: evidence; «Across the complete 87-task Terminal-Bench 2.1 suite, with two attempts per task, policies increase repeated success (pass^2) in all three GPT-5.6 tiers: 50.6% to 54.0% for Luna, 55.2% to 60.9% for Terra, and 64.4% to 73.6% for Sol.»)
- In a randomized five-arm experiment, real policies achieved 61% success on eligible tasks, while no policy achieved 39%. (confirmed by the publication itself: evidence; «To isolate the mechanism we run a randomized five-arm experiment: real policies reach 61% on eligible tasks, versus 39% without a policy»)
Primary sources:
score 69.0 out of 100 · kind: research