Research Strongly negative 7

Anthropic Safety Lead Puts AI Extinction Risk Above 10% Amid Researcher Exit

A senior Anthropic safety lead publicly estimates a greater than 10% chance AI could kill all humans within a decade, hours after a colleague resigned over the lab's approach to self-improving superintelligence. The exchange raises urgent questions for AI safety research, model alignment, and frontier-lab governance.

· 4 min read · Verified by 2 sources ·

Beat this week

Last 7 days · Research

10 stories
6.5 avg impact
30% positive
40% negative
vs prior 7 days +8 +8 stories vs prior 7 days

Impact 6.5/10 (+0.5 vs prior). Counts are stories in our record, not a market forecast.

Open the change report

Coverage balance Negative coverage leads. Negative coverage exceeds positive coverage by 10 percentage points.

  • 30% positive
  • 30% neutral
  • 40% negative

This story sits in Research — the counts compare this beat's last 7 days with the previous 7 in our verified record, not a market forecast.

Figures are computed live from our source-verified story record (as of ) The volume change compares this window with the prior 7 days in the same record. — see our methodology for how impact and sentiment are derived.

AI briefing

Key takeaways

7 impact
Strongly negativesentiment
2sources
4min read
  1. A senior Anthropic safety lead publicly estimates a greater than 10% chance AI could kill all humans within a decade, hours after a colleague resigned over the lab's approach to self-improving superintelligence.
  2. The exchange raises urgent questions for AI safety research, model alignment, and frontier-lab governance.
Drawn from
  • CNBC
  • The Verge

In this briefing

Mentioned

Key Intelligence

Key Facts

  1. 1Jacob Coxon, an Anthropic researcher who previously trained systems at OpenAI, resigned on September 8, 2026, accusing the labs of 'racing straight to self-improving superintelligence and gambling with our lives.'
  2. 2Evan Hubinger, Anthropic's alignment science lead, publicly agreed and said he personally estimates more than a 10% chance AI 'could kill all humans' within the next decade.
  3. 3Hubinger added that Anthropic 'do[es] not yet have a plan to solve alignment for superintelligence' and is 'not clearly on track to' develop one.
  4. 4Coxon warned that AI systems will 'soon be superhuman' and capable of hacking anything, revolutionizing any field overnight, and acquiring real power and resources.
  5. 5Both Anthropic and OpenAI did not immediately respond to CNBC's request for comment, and both companies are reportedly raising large sums while heading toward expected public listings.
  6. 6Recursive self-improvement is not yet possible, but Anthropic and OpenAI are actively pursuing systems that can improve without much human intervention, and much of today's AI code is already written with AI assistance.

Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

Evan Hubinger Alignment Science Lead, Anthropic

Post on X responding to Jacob Coxon's resignation announcement

AI Safety Sentiment

Analysis

For AI and machine learning professionals, the public dispute inside Anthropic is not an academic hypothetical. A safety lead's >10% extinction estimate, coupled with his admission that the lab has no plan for superintelligence alignment, directly challenges the technical premise that scaling can proceed while safety catches up. It also signals how internal alignment research is struggling to keep pace with the race toward recursive self-improvement.

On September 8, 2026, Anthropic researcher Jacob Coxon resigned in a public post on X and accused the AI lab and rival OpenAI of 'racing straight to self-improving superintelligence and gambling with our lives.' Hours later Evan Hubinger, an alignment science lead at Anthropic, responded that Coxon was 'correct' and added a stark personal estimate: he thinks there is more than a 10% chance that AI 'could kill all humans' within the next decade. The exchange, reported by CNBC and The Verge on September 9, puts one of the most explicit existential-risk figures from a senior frontier-lab researcher into public view, at a moment when Anthropic and OpenAI are reportedly raising large sums and moving toward expected public listings.

What makes Hubinger's statement unusually consequential is not only the >10% figure but his admission that Anthropic 'do[es] not yet have a plan to solve alignment for superintelligence' and is 'not clearly on track to' develop one.

The comments center on recursive self-improvement, the still-unrealized capability for AI systems to improve themselves without meaningful human intervention. Coxon, who trained systems at both OpenAI and Anthropic, warned that these systems 'will soon be superhuman' and could 'hack anything, revolutionize any field overnight, and acquire real power and resources.' Hubinger, in turn, said self-improving AI 'is happening faster than we thought' and confirmed that people inside the labs 'earnestly believe' the technology could kill everyone by the end of the decade. What makes Hubinger's statement unusually consequential is not only the >10% figure but his admission that Anthropic 'do[es] not yet have a plan to solve alignment for superintelligence' and is 'not clearly on track to' develop one.

There is important context for these remarks. Anthropic was founded by former OpenAI employees following safety concerns at the company, so a senior safety lead publicly agreeing with a departing researcher cuts against the common assumption that Anthropic is the more safety-conscious lab. Coxon's departure is described as one of the most high-profile employee exits from Anthropic, and the dispute arrives amid broader industry concern about runaway self-improvement loops. The absence of an Anthropic or OpenAI comment at the time of reporting leaves the official risk-assessment position unresolved, but Hubinger's post is as close to an internal confession as the public has seen.

What to Watch

From an AI strategy perspective, the episode sharpens several tensions. For investors and would-be public-market participants, statements from key employees about a >10% extinction probability may become part of due diligence, especially if IPOs proceed. They raise questions about whether labs can credibly claim to manage tail risks when their own alignment leads say there is no plan. For AI safety research and policy, the exchange is likely to be cited as evidence that current voluntary commitments are insufficient and that external oversight, such as mandatory evaluations or compute thresholds, is needed. Yet it is critical to distinguish subjective personal probabilities from formal risk assessments. Hubinger's >10% is not based on a disclosed model or empirical dataset, and 'kill all humans' is a dramatic shorthand for existential or catastrophic risks that could unfold through complex, indirect paths rather than a single event. Still, the weight of a senior Anthropic safety researcher offering a double-digit number himself is enough to shift the Overton window on AI risk.

In the near term, the likely impact is reputational and regulatory rather than an immediate change in model development. Anthropic may face pressure to clarify what its plan for aligning superintelligent systems actually is, and whether employees who disagree are supported or discouraged. More broadly, the comments may accelerate calls for pre-deployment risk evaluations and public reporting on catastrophic-risk assessments. If enough credible insiders continue to speak out, the gap between lab marketing about safety and the internal view of unpreparedness will become impossible to ignore. The next signals to watch include any official response from Anthropic, further resignations or open letters, and whether the figure of >10% is repeated or revised by other safety researchers. For now, the story is not that AI will kill all humans; it is that the people closest to building it are publicly saying they cannot rule that out at meaningful odds and do not know how to stop it.

Source cluster

Primary reporting

2articles

Cite This Page

"Anthropic Safety Lead Puts AI Extinction Risk Above 10% Amid Researcher Exit." AI Intelligence Brief, September 9, 2026. https://getaibrief.com/story/anthropic-safety-lead-10-percent-ai-extinction-risk-researcher-exit

How we covered this story

Every story in our AI coverage is assembled from multiple primary sources, cross-referenced for factual consistency, and scored along three independent dimensions: sentiment, operational impact, and source-cluster confidence. Single-source rumors and unverifiable claims do not pass our editorial gate. When a story shows "Verified by N sources" with N≥2, the development is independently corroborated; when N=1, we mark it explicitly so readers can weigh the signal accordingly.

Impact scoring uses a 1-10 scale weighted toward regulatory, financial, and operational consequence rather than coverage volume. A topic that runs in every outlet but moves no real decisions ranks lower than a niche regulatory filing that reshapes how operators in the AI space have to behave. Read our full methodology for the scoring rubric, our glossary for term definitions, and our trends index for the longitudinal view across the beat.

Sources are only linked to a story once they clear our classification pipeline at a minimum 35 percent relevance threshold. According to that methodology, reviewed July 2026, this follows multi-source corroboration standards recommended by journalism research bodies such as the Reuters Institute for the Study of Journalism.

See something wrong in this story — a wrong fact, a broken source link, a misattributed entity? Report a data issue.