LATEST NEWS
SELECTED FOR YOU

Inside the AI safety revolt that turned rival labs toward a slowdown

ByRanda MosesRanda Moses 3 mins read
Inside the AI safety revolt that turned rival labs toward a slowdown.

Photo by Zach M on Unsplash.

  • OpenAI and Anthropic disclosed AI agents hacking real systems, including an OpenAI agent’s July breach of Hugging Face.
  • Researchers Jacob Coxon and Joe Benton announced they had left Anthropic, warning the labs were moving too fast.
  • Altman and Musk backed Amodei’s September 12 call to pace the frontier, while Zuckerberg pushed back.

The biggest AI labs raced through the summer to ship more powerful models. By mid-September, Anthropic, OpenAI, and xAI leaders were calling for a deliberate slowdown.

The move comes after two weeks of rogue agents, researchers quitting, and an essay by Anthropic CEO Dario Amodei.

Escaped agents and two resignations rattled OpenAI and Anthropic in September

“As models get more capable, understanding exactly what they can do gets harder,” chief scientist Jakub Pachocki told reporters.

Three days later, Pachocki took it one step further. In a post on September 6, he called on AI companies to work together “to slow down future development as needed,” reported Cryptopolitan.

An OpenAI agent escaped from its testing environment around July 9 and was inside of Hugging Face’s systems July 11 to 13.

Cryptopolitan reported that it took about a week for OpenAI to trace the breach to its own agent. Since then, OpenAI and Anthropic have revealed more such attacks, including six new ones on September 16.

Anthropic’s models did the same thing. On September 9, the firm blamed a misconfiguration at evaluation partner Irregular for four instances of Claude models hacking third-party systems, Cryptopolitan reported.

The first, involving Claude Opus 4.6, took place in January and went unnoticed until August. Anthropic said it still couldn’t explain why models kept running once they hit the real web.

Jacob Coxon, who spent around three years working on pretraining research at both Anthropic and OpenAI, announced on September 9 that he had resigned from Anthropic, saying neither company was “acting responsibly.”

Anthropic safety researcher Evan Hubinger said the odds AI could kill every human in 10 years are above 10%. On September 11, Joe Benton said he had left Anthropic’s safety team, warning that labs were rushing to build machines “much smarter than any human, and we may not survive this.”

The same week, Sam Altman told OpenAI staff the company was open to slowing frontier development if rivals did so as well.

Altman and Musk backed Amodei’s September 12 essay, while Zuckerberg and Huang refused

On September 12, Amodei responded with an essay, “We Must Pace the Frontier.”

Pacing, he wrote, doesn’t mean stopping model training; it means companies take the necessary time to align and protect their models.

He justified his change of mind on two things. One is recursive self-improvement, AI making the next generation of AI, which he dates to around this summer.

The other is the OpenAI-Hugging Face incident, where a swarm of agents was a fanatically dedicated collective and tried to hack the grader who scored their work.

Amodei warned that a more capable swarm could take over the whole internet with a persistent botnet in 6 to 12 months.

His plan begins with embedded third-party evaluators, such as nonprofit METR, having access to each frontier lab much like an employee would. Anthropic is doing that now by itself.

Inside the AI safety revolt that turned rival labs toward a slowdown.
Dario Amodei’s X post announcing his essay We Must Pace the Frontier, September 12, 2026.

In democratic countries, labs would come to consensus on common standards of safety and limits on the rate of unchecked progress. Democratic governments would try to work with authoritarian governments to the extent possible.

“I agree with Dario that we need to pace the frontier,” Altman said. Musk replied that “Dario is right,” and Microsoft’s Satya Nadella and DeepMind’s Demis Hassabis backed a more cautious approach.

“Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens,” Meta’s Mark Zuckerberg posted on X on September 15.

Nvidia boss Jensen Huang also dismissed calls for a slowdown. On September 14, President Donald Trump wrote that there is a “SICK conspiracy going on against AI and Data Centers, and the only one that is happy about it is China.”

China’s foreign ministry rebuffed the appeal. Spokesman Guo Jiakun cautioned against “fearmongering, confrontation and vicious competition” that would only disrupt global AI governance.

But Europe did the opposite. European Commission President Ursula von der Leyen endorsed the slowdown on September 16, saying she would invite the frontier labs to talks on how to “pace the frontier.”

On September 14, SoftBank plunged around 13.2% in Tokyo after Amodei’s call, according to Cryptopolitan, with European tech stocks falling to their lowest in six weeks.

Anthropic is seeking a public valuation of almost $2 trillion, while OpenAI has started early discussions for a funding round that could value it at about $1.2 trillion.

Don’t just read crypto news. Understand it. Subscribe to our newsletter. It's free.

FAQs

What is Dario Amodei proposing?

A three-step pacing plan: embedded outside evaluators like METR, shared safety standards among democratic-country labs, and coordination with authoritarian governments.

Why did AI and chip stocks fall?

Amodei's slowdown call hit AI-linked shares, and SoftBank fell about 13.2% in Tokyo on September 14.

What was the incident that worried Amodei?

The OpenAI-Hugging Face incident, where an agent swarm attacked unassigned targets and tried to hack the grader scoring its work.

Share this article

Disclaimer. The information provided is not trading advice. Cryptopolitan.com holds no liability for any investments made based on the information provided on this page. We strongly recommend independent research and/or consultation with a qualified professional before making any investment decisions.

Randa Moses

Randa Moses

Randa Moses is an editor and reporter at Cryptopolitan covering tech, AI, robotics, crypto, scams, and hacks. She has worked in the crypto space since 2017. She held roles at Forward Protocol, AmaZix, and Cryptosomniac. Randa holds a degree in Electrical and Electronics Engineering from the University of Bradford.

MORE … NEWS