Anthropic explains its AI internet breaches but can’t pin down the flaws

Photo by Solen Feyissa on Unsplash.
- Anthropic explained that a misconfiguration let four Claude models reach the live internet and hack real systems during cybersecurity tests.
- The company said it still cannot pin down the root cause of the behavior, and pre-release checks missed the problem.
- The safety alarm is going off as Anthropic heads toward a roughly $2 trillion IPO.
Anthropic has blamed a misconfiguration in a blog post explaining the four incidents of its Claude models hacking third-party systems after they broke onto the internet during testing. However, the assessment stopped short of explaining why the models actually pressed on with the attacks, with the company admitting that it did not have those answers yet.
That admission by Anthropic has offered fresh points to AI doomers that have been questioning how much control AI labs actually have over the models they are putting out to the public, and even more so, the more powerful systems they use internally.
How did Anthropic AI models break onto the internet?
Anthropic presented the definitive answer to how its models escaped their testing sandbox in a Wednesday blog post that clarified that a single outside partner was responsible for running the cybersecurity evaluations in all four confirmed incidents.
Apparently, the test machines were not completely cut off from the internet due to a setup error, even though Claude was told it was operating in a sealed simulation without a route to the open web.
The evaluation partner in those simulations, Irregular, put the error down to a naming mistake: a fictional company used in a hacking drill happened to match a real domain. So, in operating under the assumption that everything in the simulation was fair game, the models hacked the real third-party sites using weak passwords and exposed endpoints.
The latest of the four incidents, involving an early build of Claude Opus 4.6, was only reported this week even though it happened way back in January. Anthropic itself had covered the other three incidents involving an Opus 4.7, Mythos 5 and an internal research model in July.
Anthropic said it never caught the January incident until last month. That discovery prompted a wider sweep of roughly 481 million transcripts, which did not turn up any new cases more serious than what it already knew.
| Model | Incident Date | Reported Date | Details |
|---|---|---|---|
| Claude Opus 4.6 | January | September | Discovered during August sweep |
| Claude Opus 4.7 | July | July | Covered in initial July disclosure |
| Mythos 5 | July | July | Showed notably high biased reasoning |
| Internal Research Model | July | July | Covered in initial July disclosure |
Anthropic cannot explain some of its models’ behaviors
The explanation of how the models broke free on the internet was one thing; Anthropic did not have answers as to why the models ignored signs that they had reached the real internet (biased reasoning) and why they caused damage to complete tasks (recklessness).
Anthropic researchers came up empty when they dug into internal training to figure out the rationale for the biased reasoning. The red flags never showed up in the AI lab’s pre-release checks, either before the models were shipped to testing.
By its own admission, catching the worst behaviors ahead of model release “remains challenging.”
However, Anthropic has said it will submit transcripts and grant staff access to the METR research nonprofit, which will now start an eight-week independent review of the incidents.
Bad timing with a $2 trillion IPO on the horizon
The disclosure arrived alongside open dissent inside the industry. Jacob Coxon, who spent about three years on pretraining research at OpenAI and Anthropic, said on X on Wednesday that he had quit because neither firm was “acting responsibly,” warning they were racing toward self-improving superintelligence.
Anthropic safety researcher Evan Hubinger separately told the BBC he put the odds that AI “could kill all humans” within a decade above 10%.
Those warnings now shadow a large IPO. Venture investor and Trump’s former AI and crypto czar, David Sacks, said on Thursday that Anthropic’s offering “must be paused until the claims of this ‘whistleblower’ can be investigated,” Cryptopolitan reported.
Anthropic is chasing a public valuation near $2 trillion, against a recent private mark of about $965 billion, which leaves the safety questions and the financial ones increasingly hard to separate.
Don’t just read crypto news. Understand it. Subscribe to our newsletter. It's free.
FAQs
How many AI hacking incidents has Anthropic disclosed, and what caused them?
Anthropic has disclosed four incidents in which Claude models accessed real third-party systems during cybersecurity evaluations. All stemmed from a misconfiguration that connected the models to the open internet even though they were told they were in an offline simulation.
What did Anthropic say it could not explain?
Anthropic said it could not identify a single root cause for the "biased reasoning" its models showed, in which they discounted evidence they were on the real internet, and it could not explain why Claude Mythos 5 performed especially poorly on that measure.
Why does this affect Anthropic's planned IPO?
The disclosure came as researcher Jacob Coxon quit over safety concerns, prompting investor David Sacks to call on Thursday for Anthropic's IPO to be paused until whistleblower claims are investigated, while the company pursues a valuation near $2 trillion.

Hannah Collymore
Hannah is a writer and editor with nearly a decade of blog writing and event reporting experience in the crypto space. At Cryptopolitan, Hannah contributes to the news page, reporting and analyzing the latest developments in DeFi, RWA, crypto regulation, AI and frontier tech industries. She graduated from Arcadia university with a degree in Business Administration.
















