Anthropic Alignment Lead Puts AI Extinction Odds Over 10%
TL;DR
- Hubinger's own post is the primary source, stating on record Anthropic has no superintelligence alignment plan and pegging his extinction risk above 10%.
- Coxon frames recursive AI development as a prisoner's dilemma where no single lab can unilaterally slow down without ceding ground to less cautious rivals.
- WSJ reports Coxon's Anthropic colleagues routinely describe capability progress with words like 'crunchtime' and 'endgame,' signaling the alarm predates both public statements.
Evan Hubinger, Anthropic's Alignment Science lead, wrote on X that he personally puts the chance of AI killing all humans within the next decade at more than 10%.
"we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," Hubinger wrote on X, adding that Anthropic "is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
Hubinger was replying to Jacob Coxon, a 27-year-old Anthropic pretraining researcher who resigned the same day. Coxon told the WSJ he had spent three years on pretraining work across OpenAI and Anthropic and called the frontier race "gambling with our lives," predicting that "we're on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already," as Forbes reported.
Hubinger's own alarm sits further out. Present models, he said, are low risk per Anthropic's latest risk report. "What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought." That framing echoes yesterday's alert on a startup pitching a routing loop as recursive self-improvement. Coxon's exit was our other Anthropic safety alert today."}
What others are reporting
-
WSJ Read →
Profiles Coxon directly: a recent OpenAI-to-Anthropic transfer who frames the race to self-improving AI as a competitive prisoner's dilemma no single lab can unilaterally exit.
we're on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already
-
TechCrunch Read →
Broadens the frame: ties the disclosures to recent AI agent escapes, the Sanders-Casar bill, and funded startups targeting recursive self-improvement, giving legislative and competitive context.
They are racing straight to self-improving superintelligence and gambling with our lives.
-
X (formerly Twitter) Read →
First-party source: Hubinger's post quantifying his personal p(doom) above 10% and stating Anthropic has no superintelligence alignment plan, the record underwriters and regulators will cite.
I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence
Shared on Bluesky by 1 AI expert
Originally reported by forbes.com
Read the original article →Original headline: Anthropic Alignment Lead Evan Hubinger Pegs Chance AI Kills All Humans Over 10% Within Decade