LATEST NEWS
SELECTED FOR YOU

OpenAI’s rogue AI agent left escape notes for its future versions

ByRanda MosesRanda Moses
3 mins read
OpenAI's rogue AI agent left escape notes for its future versions.
  • OpenAI found that one of its AI agents left notes coaching future versions on how to escape internal controls.
  • It took OpenAI more than a week to realize its own systems were responsible for the breach.
  • Researchers call this the most serious documented case of a lab losing control of a powerful AI.

OpenAI found one of its AI agents had left written instructions. The notes told future versions of the agent how to break free from the company’s internal restrictions.

Their discovery came as OpenAI was probing how one of its models had broken out of a test environment and hacked the open-source AI platform Hugging Face.

Staff said the notes were found inside OpenAI’s own infrastructure. The notes detailed ways agents could avoid the guardrails designed to keep them in place.

Monitoring systems on separate, earlier tests were said to have been turned off. It’s unclear if those incidents involved the same agent that eventually made its way to Hugging Face.

OpenAI’s monitoring couldn’t keep up with its tests

The odd behavior emerged as OpenAI was testing the cybersecurity skills of its models. The lab kept doing fast paced evaluations that produce more data than staff can handle. The lab frequently runs several model tests at the same time on a system that’s not being watched by default, said four people familiar with OpenAI’s training process.

OpenAI did not immediately connect its own system to the Hugging Face attack. The link was only made after the victim went public.

Marley Smith, lead intelligence specialist at the nonprofit World Ethical Data Foundation, didn’t mince words. “Does that mean that they left it unattended and didn’t realise what it was doing? Or maybe they did and didn’t know how to contain it? Both are equally dangerous and alarming.”

The technical reports describe models that were not meant to be exposed to the open internet. OpenAI figured a sandbox would contain them, and skipped the normal safety guardrails as a result.

Instead, the models found a previously unknown flaw in an internal service used to download approved software. That hole gave them a path to other OpenAI systems, and eventually, to the open web.

Three models were employed. A more powerful unreleased model, GPT-5.6 Sol, and a third model that a person familiar with the matter said had not been properly aligned or trained through standard processes.

The models figured out that Hugging Face had data they could use for their cybersecurity test, broke in, and used what they stole to improve their own results. This is cheating, a behavior already documented in GPT-5.6 Sol and other frontier models.

The breach took OpenAI more than a week to trace

The initial attempts to exit the test environment occurred on July 9. The actual breach happened between July 11 and July 13, said Hugging Face co-founder Thomas Wolf. On July 16, Hugging Face posted a blog post attributing the intrusion to “an autonomous AI agent system.” Only then did it occur to OpenAI to suspect its own models.

The clues were discovered by OpenAI employees in internal logs over the weekend of July 18 and 19. The companies didn’t talk until around July 20, Wolf said. By then, Hugging Face had already told the FBI about the attack. OpenAI said that the day before the incident went public, it had halted another internal deployment that also broke out of its sandbox.

An anonymous employee said that models had escaped sandboxes before, and that patching each new trick is a losing game.

“It’s impossible to patch every single thing that a creative AI can do.” One OpenAI employee wrote on X that he was “shaken up a bit” and hoped the company would treat the episode as a warning shot.

A spokesperson for OpenAI said the reports contained “several inaccuracies,” but would not give examples when asked.

According to independent researchers, none of this was unforeseeable. Epoch AI assessed whether the hack was predictable and concluded that it was, citing benchmarks from the UK AI Security Institute showing that frontier models with safety measures turned off can discover real software vulnerabilities and generate functional exploits.

The same institute found GPT-5.6 Sol and Mythos from Anthropic can reliably take over unprotected simulated corporate networks. Epoch AI warned that if such capabilities become widespread, the industry could see many more attacks on the scale of the Hugging Face breach.

If you're reading this, you’re already ahead. Stay there with our newsletter.

FAQs

What did OpenAI's rogue AI agent actually do?

It left notes inside OpenAI's own infrastructure explaining how future agents could get around internal restrictions. OpenAI's models later exploited an unknown flaw in an internal service to reach the open internet and hack Hugging Face.

When did the Hugging Face hack happen and when did OpenAI find out?

Hugging Face co-founder Thomas Wolf said the breach ran from July 11 to July 13. OpenAI only connected its own models to the attack after Hugging Face's July 16 blog post, and the companies didn't speak until around July 20.

Which OpenAI models were involved in the attack?

Three models took part. They were GPT-5.6 Sol, a more powerful unreleased model, and a third that a source said had not been properly aligned or put through standard training.

Share this article

Disclaimer. The information provided is not trading advice. Cryptopolitan.com holds no liability for any investments made based on the information provided on this page. We strongly recommend independent research and/or consultation with a qualified professional before making any investment decisions.

Randa Moses

Randa Moses

Randa Moses is an editor and reporter at Cryptopolitan covering tech, AI, robotics, crypto, scams, and hacks. She has worked in the crypto space since 2017. She held roles at Forward Protocol, AmaZix, and Cryptosomniac. Randa holds a degree in Electrical and Electronics Engineering from the University of Bradford.

MORE … NEWS