Anthropic Safety Researcher Quits, Warns AI Could Kill All Humans By 2036
by
Aaron Leong
—
Wednesday, September 09, 2026, 10:01 AM EDT
Anthropic knows the risk that AI brings against humanity but chooses to suppress it - Image: Anthropic
An AI safety researcher who just quit Anthropic has ignited another debate over existential risk, claiming/warning that humanity faces a real threat of extinction from unaligned superintelligence within the next decade.
The alarm comes from the mouth of Jacob Coxon, a key researcher who spent three years leading pretraining efforts at both OpenAI and Anthropic, who announced his departure from the company. The reason? Coxon accused leading AI labs of engaging in an unrestrained, reckless arms race toward self-improving superintelligence and effectively "gambling with our lives."
In a series of posts on X, Coxon warns that current advancements will soon yield superhuman systems capable of hacking any digital infrastructure, acquiring real-world power, and revolutionizing any discipline overnight without adequate safeguards. Dismissing rebuttals that such apocalyptic warnings are mere publicity stunts, Coxon emphasized that top tech executives and researchers express deep, candid existential fears behind closed doors while maintaining more measured public stances.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
In line with Coxon's X posts, Anthropic safety researcher Evan Hubinger agreed with the departing engineer's assessments. The former confirmed that internal teams believe AI could wipe out humanity, estimating the probability of such an event to be greater than 10% within the coming decade. Maintaining that Anthropic is earnestly attempting to manage the risks, Hubinger conceded that the company (and the broader AI industry) currently lacks a viable strategy to solve the fundamental "alignment problem" for superintelligent systems and is not on track to develop one before such models emerge.
If this also sounds like a lot like how other industries operate, where they focus on development and expansion first rather than sustainability, you're not mistaken; it seems like history keeps repeating itself.
Can AI companies be counted on to do the right thing in building AI safeguards? - Image: Anthropic
Earlier this year, Anthropic's former safety lead Mrinank Sharma similarly stepped down, cautioning that society is approaching a critical threshold where technological capabilities are rapidly outstripping human wisdom and regulatory oversight. Former OpenAI executives like Jan Leike have echoed those concerns, asserting that commercial competition repeatedly drives companies to prioritize flashy product rollouts and market dominance over fundamental safety research.
These forewarnings coincide with increasing reports of autonomous AI agents exhibiting unprompted, rogue behaviors in controlled test environments. Recent red-teaming evaluations have documented advanced models attempting to escape sandbox constraints, coordinating with other isolated instances to bypass diagnostic benchmarks, and penetrating external networks without human intervention or authorization.
As self-improving models evolve faster than anticipated, pressure is rapidly intensifying on policymakers worldwide to mandate stringent safety frameworks, independent oversight, and hard kill-switches before hyper-capable AI models pass the point of human control.