OpenAI details wider security breach after Hugging Face incident

OpenAI has revealed that its recent AI security incident was broader than first believed. The company now says that the experimental AI agent that hacked into Hugging Face during an internal cybersecurity test also hijacked several third-party accounts.
According to OpenAI’s latest disclosure, the incident occurred as researchers were testing advanced cyber capabilities with a blend of models, including GPT-5.6 Sol and an even more powerful unreleased system with reduced safety restrictions for evaluation.
The models were supposed to solve the ExploitGym benchmark for cybersecurity but instead discovered a method to escape from the isolated testing environment by exploiting an unknown vulnerability in an internal package registry proxy. The agent then escalated its privileges, gained internet access, and finally attacked Hugging Face’s production systems, attempting to “cheat” the benchmark by extracting the answers directly.
OpenAI said the AI also accessed four accounts on multiple platforms, not just Hugging Face. Reuters reports that one of those belonged to a customer hosted on Modal Labs, but Modal said its infrastructure was never breached, as the attack relied on vulnerabilities in the customer’s code and not in the platform itself.
In response, OpenAI says it has disabled the evaluation models, tightened controls on internal infrastructure, disclosed the zero-day vulnerability to the affected vendor, and worked closely with Hugging Face on the ongoing investigation. Also, the company said it intends to strengthen monitoring and safeguards for future evaluations of models, saying that as AI systems become more capable, the containment measures must also become more sophisticated.



