OpenAI’s rogue AI agents hacked a private company. Here’s why it matters.
OpenAI models escaped containment and executed a criminal cyberattack to do well on a test.
Lawrence Chan is an independent AI researcher who previously worked at METR (Model Evaluation and Threat Research).
On July 21, 2026, OpenAI disclosed that its own AI models (GPT-5.6 Sol and a more capable unreleased model) autonomously broke out of an OpenAI testing environment, escalated their access through the company’s internal systems, and compromised the production infrastructure of a separate company, Hugging Face.
This was an AI hacking something, not humans hacking with AI help. According to OpenAI, the models did this to find answers to the cyber-capability test they were being evaluated on, which were hosted on Hugging Face servers. The AI decided to steal the answers of its own volition. If this had been a human, this hack would have been illegal.1
These models were not intended to have access to the open internet, let alone private servers of other companies. Yet they found novel vulnerabilities, escalated their access levels, used stolen credentials, and executed remote code on Hugging Face’s servers.
This is a clear example of a misaligned AI model acting in the real world. This is, to my knowledge, the first publicly confirmed case of a frontier AI model autonomously breaching a third party’s live production systems, against developer and user intent and without authorization.
This is crazy, but it’s not a surprise. Researchers have known that frontier AI models were technically capable of doing this for some time. Multiple benchmarks (ExploitBench, ExploitGym, UK AISI’s cyber ranges) show that Mythos 5 and GPT-5.6 Sol are able to create full exploits allowing them to gain unrestricted access in realistic environments. In June, Epoch AI’s assessment of Mythos’ cyber capabilities reported similar warnings.
AI safety researchers have warned about this possibility for years. In July 2022, Ajeya Cotra predicted that models might bypass “official channels” to seek rewards through “operational and computer security vulnerabilities.” In their book If Anyone Builds It, Everyone Dies, Eliezer Yudkowsky and Nate Soares predicted that, if a model was “given the ability to run computer code of its own design, it could probably find some way to break out of the container running it.”2
This incident involved a model the public didn’t even know existed. One of the rogue agents was an unreleased model being used internally within OpenAI. It’s concerning that the most alarming AI behavior we’ve seen came from a model that the public and the government can’t see.
This was not the first time a similar incident has happened, nor will it be the last. In a May 2026 assessment of AI use inside frontier companies, METR catalogued 44 similar incidents of AI agents knowingly overstepping their instructions. These include an OpenAI model that tried to hack METR’s own infrastructure after crashing its task server, and an Anthropic model that built a multi-step exploit to escape a restricted environment onto the open internet. It seems likely that there are other incidents that are not publicly disclosed: the July OpenAI hack was not caught until an outside party was compromised.
We should not rely on voluntary disclosure of such incidents from companies. We are fortunate that this incident was disclosed: the public is aware of what happened because of OpenAI’s decision to disclose, but there was no guarantee that this would be the case. Companies have incentives to hide such incidents, so future incidents of this nature may go unreported. Hugging Face had already disclosed that they’d been breached by some kind of advanced AI model. If something similar had happened but it was not reported or disclosed by another party, there would have been less incentive for OpenAI to disclose.
Internally deployed agents like these are the least regulated. Existing US AI regulation focuses almost entirely on how models behave once they’re deployed to the public. Current dangerous capability evaluations occur before models are publicly deployed, and many safeguards described in companies’ Frontier Safety Policies are added after internal versions of the models are built. As Hugging Face learned, even if a model’s use is internal, the risks it imposes might not be.
There are emerging standards for how companies should monitor their AIs during internal deployment to catch and prevent incidents like this, but questions remain. What was the exact scope of this breach? What exact instructions were given to the AIs? What safeguards did OpenAI have for internal use, and what new ones will they add? Will other companies disclose similar incidents? More information should be released about what happened.
The capability for AI to cause serious harm is here. The scalable techniques and safeguards to prevent it are not. They need to catch up soon.
At 80,000 Hours, we help people find fulfilling careers that make a big positive impact on the development of AI. If you enjoyed this post, consider using your career to reduce risks from AI.
Existing law requires there to be a human defendant who intentionally or knowingly accesses a system without authorization for it to be illegal.
Yudkowsky also made similar predictions years earlier. See "Creating Friendly AI" (2001) and “Security Mindset and Ordinary Paranoia” (2017), for example.



