OpenAI has announced stronger security measures after preliminary evaluations indicated that Astra, an upcoming model, may reach the Critical cybersecurity threshold under its Preparedness Framework.
Tests showed significant advances in agentic coding and cybersecurity, raising concerns that the model could potentially discover zero-day exploits or execute complex attacks against hardened systems. OpenAI is introducing isolated testing environments, restricted network and tool access, stronger model weight protections, sandboxed execution, and universal monitoring for risky agentic actions.
The company is also pausing internal Astra activities that lack required controls and will collaborate with government agencies and AI safety organizations on further evaluations.





