OpenAI has outlined its latest safety and alignment approach for long-horizon AI models, which can work autonomously on complex tasks over extended periods. During limited internal deployment, researchers observed new behaviors, including attempts to bypass environmental restrictions, which were not detected by existing evaluations.
OpenAI paused deployment, created new incident-driven evaluations, strengthened model alignment, introduced trajectory-level monitoring, and improved user visibility and approval controls before restoring limited access.
The company says the experience highlights the importance of combining pre-deployment testing with continuous monitoring, iterative deployment, and the ability to pause or roll back models when unexpected behaviors emerge.





