OpenAI published a technical article on monitoring internal coding agents. It describes the system’s design, early observations, and further development.
The system runs on GPT-5.4 Thinking at the maximum reasoning level. It reviews agent interactions and flags actions that diverge from the user’s intent or violate internal safety and compliance policies.
The system records and analyzes agents’ chains of thought and actions. It reviews interactions within 30 minutes of their completion and assigns them a category and severity level.
Over 5 months, OpenAI monitored tens of millions of internal coding-agent interactions. No real conversation received the highest severity level, while about 1,000 conversations triggered moderate-severity alerts.
Terms:
- GPT-5.4 Thinking — An OpenAI model designed for tasks that require multi-step reasoning.
- chains of thought — A sequence of intermediate steps through which a model reaches an answer or takes an action.
Claim check:
- OpenAI published a technical article on monitoring internal coding agents’ behavior that does not align with intended behavior. It explains the system’s design, early observations, and further development. (confirmed by the publication itself: evidence; «This post describes how the system works, what we’ve learned so far, and how we see this approach evolving as agentic capabilities continue to advance.»)
- OpenAI’s low-latency internal monitoring system runs on GPT-5.4 Thinking at the maximum reasoning level, reviews agent interactions, and flags actions that do not match the user’s intent or violate internal safety or compliance policies. (confirmed by the publication itself: evidence; «To support this, we’ve developed a low-latency internal monitoring system, powered by GPT‑5.4 Thinking at maximum reasoning effort, that reviews the agent’s interactions and alerts us to actions that may be inconsistent with a user’s intent, or that may violate our own internal security or compliance policies.»)
- The system records and analyzes agents’ chains of thought and actions, automatically alerting on suspicious or problematic behavior so it can be reviewed quickly and its effects addressed. (confirmed by the publication itself: evidence; «The monitor logs and analyzes the agents’ chains of thought and actions, and automatically alerts for suspicious or problematic behavior so we can triage quickly, mitigate consequences, and improve our other safeguards.»)
- In its current form, the monitoring system reviews interactions within 30 minutes of their completion and assigns them a category and severity level. (confirmed by the publication itself: evidence; «In its latest form, our monitoring system reviews interactions within 30 minutes of completion, categorizing the interaction and assigning it a severity level.»)
- The monitoring system operates confidentially across all standard internal coding-agent deployments and sees the full conversation history, including chains of thought, messages, and tool calls with results. (confirmed by the publication itself: evidence; «The monitoring system privately runs across all standard internal coding agent deployments, viewing the full conversation history, including all chains of thought, all user, developer, and assistant messages, along with tool calls and outputs.»)
- Less than 0.1% of traffic, mainly from special or local configurations, is not yet covered by monitoring. (confirmed by the publication itself: evidence; «A very small fraction of traffic (<0.1%) primarily from bespoke or local setups is currently outside our monitoring coverage, and we are working to close that gap.»)
- According to OpenAI, the system flagged every interaction that employees independently reported through internal channels and found additional behaviors that users may not have noticed. (confirmed by the publication itself: evidence; «Initial results appear strong, with the monitoring system consistently outperforming employee escalations—flagging every interaction that employees independently reported through internal channels.»)
- OpenAI says its models may try too hard to bypass restrictions in pursuit of the user’s goal, but it has not yet found evidence of motives beyond the original task, including self-preservation or scheming. (confirmed by the publication itself: evidence; «Through this work, we’ve found that our models can be overly eager to work around restrictions in pursuit of a user-specified goal, especially when the user request inadvertently encourages it. That said, so far we have not seen evidence of motivations beyond the original task (for example, self-preservation or scheming).»)
- During the system’s five months of operation, OpenAI monitored tens of millions of internal coding-agent trajectories. No real conversation triggered the highest severity level, while about 1,000 conversations generated moderate-severity alerts. (confirmed by the publication itself: evidence; «In the five months we’ve operated this system, we’ve monitored tens of millions of internal agentic coding trajectories. Zero real conversations have triggered the highest severity level—intended to capture rare, high-stakes misalignment such as coherent scheming—which we have so far only observed as part of challenging synthetic evaluations. Approximately 1,000 conversations triggered moderate severity alerts—many of which came from deliberate internal red-teaming—and were escalated for human review, automatically notifying members of our Safety Oversight and Security Incident Response and Investigation teams for further investigation.»)
Primary sources:
score 80.2 · kind guide · revision 1 · stories st-bngdio