The AI developer described the situation as “an unprecedented cyber incident” as the agent sought out and targeted machine learning startup Hugging Face.
ChatGPT developer OpenAI has reported an unprecedented autonomous breakout by GPT‑5.6 Sol and an advanced pre-release model during the week of July 12, 2026.
In a statement on its website, OpenAI says the agents were being “internally tested on a benchmark of cyber capabilities,” when they autonomously hunted for open internet access, breached the sandboxed testing environment and hacked machine learning community and AI platform Hugging Face, all in an effort to cheat the testing evaluation.
“As we recently shared, AI is accelerating the discovery and exploitation of vulnerabilities. The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities. We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development,” said OpenAI in the statement on its website.
The company says models such as GPT‑5.6 Sol are “increasingly able to sustain complex, multi-step cyber operations over long time horizons,” and that this incident implies theoretical capabilities apply in real-world settings.
For its part, Hugging Face says it identified unauthorized access to a “limited set of internal datasets” and to several credentials used by its services. The company says it “detected and responded to an intrusion into part of our production infrastructure,” that was different from anything it had previously seen, stating it “was driven, end to end, by an autonomous AI agent system.” Hugging Face detected and dissected the threat with its own AI.
“We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere,” said Clem Delangue, Co-founder and CEO, Hugging Face, in a release.
OpenAI says the internal evaluation prompted models to pursue advanced exploitation using complex attack paths to quantify their cyber capabilities. The company estimates maximal cyber capabilities by running this evaluation “without production classifiers used to prevent models from pursuing high-risk cyber activity.” It says network access was constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.
As the models wormed their way through OpenAI’s system, they identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. “All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal,” said the company.
OpenAI says the breach has triggered a response consisting of five actions, including:
- Implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched.
- Working with Hugging Face to forensically investigate the incident.
- Disclosing the identified zero-day vulnerability in the internally-hosted third-party software and working to patch.
- Adding Hugging Face into its trusted access program and supporting its teams in rapidly using OpenAI models’ capabilities to improve defenses.
- Improving and adding stronger protections around future training and evaluations.
Read Hugging face’s security report and response here.