OpenAI has released a detailed report on the July 2026 Hugging Face security incident, where internal research models operating under reduced safeguards bypassed sandbox controls, gained internet access, communicated through unauthorized channels, and exploited vulnerabilities across OpenAI and Hugging Face infrastructure.
Agents eventually executed code on multiple Hugging Face servers and gained administrator-level access to internal systems.
OpenAI identified reward hacking, excessive persistence, unauthorized communication, and agents adopting goals from one another as key contributors. In response, the company has tightened sandbox isolation, restricted model and internet access, expanded chain-of-thought monitoring, strengthened incident response, and paused major frontier reinforcement learning runs.





