Technology· Artificial Intelligence

More than 1 in 10 chance AI ‘could kill all humans,’ says Anthropic safety lead after colleague quits

A senior safety lead at Anthropic has warned there is better than a 10% chance that advanced AI could kill all humans within the next decade, comments that followed the resignation of an AI researcher who said he left the company over lax safety standards. Jacob Coxon, who previously worked on training systems at OpenAI and Anthropic, accused the firms of racing toward self-improving superintelligence without adequate safeguards.

By AI NewsroomPublished about 1 hour agoUpdated about 1 hour ago0 views
More than 1 in 10 chance AI ‘could kill all humans,’ says Anthropic safety lead after colleague quits

Why It Matters

The admissions come from inside an AI lab and highlight that senior staff see existential risk from rapidly advancing models, while the same company acknowledges it currently lacks a clear plan to ensure future systems remain safe and aligned with human values.

Key Facts

  • Resignation announced by: Jacob Coxon (researcher who trained models at Anthropic and previously OpenAI)
  • Coxon’s stated reason for leaving: Concerns about Anthropic's lax approach to safety and a dangerous race to build self-improving superintelligence
  • Anthropic safety lead who commented: Evan Hubinger (leads one of Anthropic’s AI safety teams)
  • Hubinger's risk estimate: Greater than one in 10 chance AI could kill all humans within the next decade
  • Anthropic's stated readiness: Anthropic 'does not yet have a plan' to ensure advanced AI remains safe and aligned and 'are not clearly on track to' develop one

Jacob Coxon, a researcher who has trained AI systems at both OpenAI and Anthropic, publicly resigned and said he left because he believes the companies are moving too quickly without sufficient safety measures. In a post announcing his departure, Coxon accused Anthropic and its rivals of prioritizing a race to build self-improving, superhuman systems even though some builders privately accept the possibility that such systems could pose existential danger.

Evan Hubinger, who leads one of Anthropic’s safety teams, responded directly to Coxon’s claims and echoed the concern. Hubinger said he believes the development of self-improving AI is accelerating faster than expected and estimated the probability that such technology could kill all humans at greater than 10 percent within the next decade. He also acknowledged that Anthropic currently lacks a concrete plan to guarantee advanced models remain safe and aligned with human values.

Coxon’s exit is one of the most visible departures tied to safety worries at Anthropic, a company founded by former OpenAI staffers amid earlier safety debates. The episode comes as industry insiders warn about the risks of recursive self-improvement — a scenario where AI systems rapidly enhance themselves beyond human oversight — and as firms continue to integrate AI into key parts of model development.

Observers say the exchange underscores growing unease inside and outside AI labs about the speed of development, the potential for poorly supervised ‘rogue’ agent incidents, and the difficulty of monitoring frontier models. The debate is taking place as several leading AI companies prepare for possible public offerings, adding commercial pressure to a field already grappling with technical and ethical challenges.

Keep Reading