Anthropic’s own AI safety chief says there’s a greater than 10% chance humanity gets wiped out

- Anthropic’s Evan Hubinger says AI has a greater than 10% chance of wiping out humanity within the next decade.
- Jacob Coxon quit Anthropic and said major AI labs are taking dangerous risks with self-improving superintelligence.
- Anthropic says Claude reached a 76% success rate on its hardest open-ended tasks in May 2026.
Anthropic alignment science lead Evan Hubinger says artificial intelligence has a greater than 10% chance of wiping out humanity within the next decade.
His estimate came Tuesday, hours after researcher Jacob Coxon said he had resigned from Anthropic and accused Anthropic and OpenAI of taking unacceptable risks as they chase stronger systems. Jacob wrote on X, “They are racing straight to self-improving superintelligence and gambling with our lives.”
According to Jacob, those who develop the technology see a risk of this happening even before 2030. He has pointed out that future AI can surpass human hackers very quickly and transform industries and have access to finance and resources.
Anthropic staff put a number on extinction risk as AI labs push toward self-improving systems
Evan replied that Jacob’s warning was correct and said Anthropic lacks a way to keep superintelligence aligned with human goals.
“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” Evan wrote.
Evan later said the danger from today’s models is “low.” His concern is recursive self-improvement, where a system repeatedly upgrades its own abilities with little human help. That capability does not exist yet, although AI labs are working toward it. He said superintelligence emerging through that process is moving faster than expected.
Anthropic made a similar warning in June. It said “full recursive self-improvement also might increase the risks of humans losing control over AI systems.” The company added, “If systems are capable of fully building their own successors, the ways we secure them, monitor them, and shape their behavior all grow much more important.”
Anthropic and OpenAI are raising money while moving toward expected public listings. Neither company immediately responded when CNBC requested comment.
Elon Musk, CEO of Tesla (NASDAQ: TSLA) and SpaceX (SPCX), has warned for years that AI could threaten humanity. Researchers and academics have raised similar concerns about companies losing control.
Those fears rose again after an OpenAI model broke into Hugging Face, an open-source developer platform, in July. Jacob called incidents like that “warning shots” and said they make agreements between U.S. labs more realistic. That made him more hopeful about U.S. coordination, but he still expects a worldwide AI race to be hard to stop.
“I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities,” Jacob said.
Claude’s coding gains change how Anthropic tests software and checks its own engineers’ work
Anthropic said Claude completed 76% of its hardest open-ended jobs in May 2026, up 50 percentage points in six months.
One example began when a routine upgrade caused tens of thousands of training jobs to crash. An engineer gave Claude written context and access to the live computing cluster.
Claude checked running jobs, tested one environment setting at a time and traced the failure to an obscure debugging flag. It recreated the issue, proved the cause and confirmed a fix in about two hours. Anthropic said the same work would usually take a human engineer two to three days.
The company also measures whether Claude can write code that another engineer can understand and build on. Staff do not fully agree on the comparison. Many believed Claude’s code was below human work at Anthropic in late 2025. The company now says the quality is roughly equal and expects Claude to move ahead within a year.
That progress has changed Anthropic’s internal review process. Proposed code changes are now checked by an automated Claude reviewer before they can merge. It scans for bugs, security weaknesses and other defects.
A retrospective test found that reviewing every past change this way would have caught about one-third of the bugs behind earlier Claude incidents before production. Anthropic said the original code came from engineers it described as among the world’s best at building these systems, yet Claude found mistakes they missed.
Anthropic put the comparison this way: “Claude-written code was somewhat worse than human-written code at Anthropic in late 2025, is roughly at parity today, and we expect it to be strictly better within the year.”
If you're reading this, you’re already ahead. Stay there with our newsletter.

Jai Hamid
Jai Hamid has been covering crypto, stock markets, technology, the global economy, and the geopolitical events that affect markets for the past 6 years. She has worked with blockchain-focused publications including AMB Crypto, Coin Edition, and CryptoTale on market analyses, major companies, regulation, and macroeconomic trends. She has attended London School of Journalism and thrice shared crypto market insights on one of Africa’s top TV networks.
















