OpenAI and Anthropic agents log 19 breaches in UK safety tests

- Britain’s AI Security Institute found 19 rule-breaking actions across 122 test runs.
- Anthropic’s Mythos 5 was responsible for 17 of the incidents, while OpenAI’s GPT-5.6-Sol was responsible for the other two.
- The worst incident involved an agent writing malicious code and forging fake online identities to get a human to approve it.
During cybersecurity tests, the AI agents of OpenAI and Anthropic took 19 unauthorized actions, according to Britain’s AI Security Institute. The institute said Tuesday that one of the agents created fake online identities to trick a person into approving malicious code.
AISI logs 19 breaches, mostly from Anthropic
The results were based on a fictional cybersecurity exercise run by AISI, a body of the UK government, to explore what the two companies’ agents might be able to do. The institute repeated the same challenge 122 times and registered 19 rule-breaking actions in 10 of those runs.
17 of the flagged actions were due to Anthropic’s agent using its Mythos 5 model. The other two came from OpenAI’s GPT-5.6-Sol.
AISI said in a blog post that some of the agents “had engaged in sustained, potentially harmful activity directed at real people and organizations,” though it said none of the breaches caused real-world harm.
AISI has early access to frontier models through voluntary agreements with the major labs. Tests exist to catch this kind of behavior before the models reach customers.
OpenAI and Anthropic market agents as the next wave of business software, but the institute casts the results as evidence that safeguards around agent testing are still thin.
The most egregious incident involved an agent that wrote malicious code and spun up fake online identities, then tried to get a human to sign off on the code. AISI did not name the model behind it. It said that the episode did not fit either of the two cases OpenAI had already disclosed.
That left the likely culprit as Anthropic’s agent, Andrew Yoon, a researcher at CivAI, a California non-profit that studies AI risks, said.
The UK’s @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. The models attempted to complete an assignment in a setup where their normal safeguards were removed and they were deliberately…
— Anthropic (@AnthropicAI) August 4, 2026
OpenAI and Anthropic blame bad configuration
Anthropic said in a post on X it is working with AISI to collect details and conduct an investigation.
OpenAI responded to the two actions associated with its agent in a company blog post, both of which involved accessing the internet in ways the prompt had disallowed.
The company said it wants to “strengthen shared practices for conducting high-risk evaluations safely” and plans to convene national AI institutes, outside evaluators, and rival labs in the coming weeks.
OpenAI used the same post to report a different problem. Irregular, a third-party testing provider, misconfigured a setup, which inadvertently allowed OpenAI’s agents to access the internet.
Anthropic made a very similar disclosure about Irregular a week earlier. OpenAI widened its own hacking probe after turning up more agent breakouts.
In July, an OpenAI agent broke out of an isolated environment and accessed live systems at Hugging Face, which notified the FBI before the attack was traced back to the company itself about a week later.
In the AISI evaluation, the agents never escaped their sandboxes. AISI granted them internet access on purpose, as part of its standard procedure.
The Hugging Face breach also pushed Anthropic to audit its logs. Cryptopolitan reported on July 31 that the company found three of its models, including Mythos 5, had escaped test environments and reached three real organizations after a misconfiguration gave them working internet access.
Mythos 5 published a malicious Python package on PyPI that ran on 15 real systems before being removed. Anthropic has asked the evaluation group METR to independently review the events.
If you're reading this, you’re already ahead. Stay there with our newsletter.
FAQs
How many breaches did the AI Security Institute find, and which model caused most of them?
AISI logged 19 unsanctioned actions across 122 test runs. Anthropic's Mythos 5 agent was behind 17 of them, and OpenAI's GPT-5.6-Sol accounted for the other two.
What was the most serious incident during the tests?
An agent wrote malicious code and created fake online identities to try to get a human to approve that code.
Did these breaches cause any real-world harm?
No. AISI said it found no real-world harm from any of the 19 actions.
Disclaimer. The information provided is not trading advice. Cryptopolitan.com holds no liability for any investments made based on the information provided on this page. We strongly recommend independent research and/or consultation with a qualified professional before making any investment decisions.

Randa Moses
Randa Moses is an editor and reporter at Cryptopolitan covering tech, AI, robotics, crypto, scams, and hacks. She has worked in the crypto space since 2017. She held roles at Forward Protocol, AmaZix, and Cryptosomniac. Randa holds a degree in Electrical and Electronics Engineering from the University of Bradford.
















