Anthropic Safety Researcher Quits, Warns AI Could Kill All Humans By 2036

Anthropic knows the risk that AI brings against humanity but chooses to supress it
Anthropic knows the risk that AI brings against humanity but chooses to suppress it - Image: Anthropic
An AI safety researcher who just quit Anthropic has ignited another debate over existential risk, claiming/warning that humanity faces a real threat of extinction from unaligned superintelligence within the next decade.

The alarm comes from the mouth of Jacob Coxon, a key researcher who spent three years leading pretraining efforts at both OpenAI and Anthropic, who announced his departure from the company. The reason? Coxon accused leading AI labs of engaging in an unrestrained, reckless arms race toward self-improving superintelligence and effectively "gambling with our lives."

In a series of posts on X, Coxon warns that current advancements will soon yield superhuman systems capable of hacking any digital infrastructure, acquiring real-world power, and revolutionizing any discipline overnight without adequate safeguards. Dismissing rebuttals that such apocalyptic warnings are mere publicity stunts, Coxon emphasized that top tech executives and researchers express deep, candid existential fears behind closed doors while maintaining more measured public stances.
In line with Coxon's X posts, Anthropic safety researcher Evan Hubinger agreed with the departing engineer's assessments. The former confirmed that internal teams believe AI could wipe out humanity, estimating the probability of such an event to be greater than 10% within the coming decade. Maintaining that Anthropic is earnestly attempting to manage the risks, Hubinger conceded that the company (and the broader AI industry) currently lacks a viable strategy to solve the fundamental "alignment problem" for superintelligent systems and is not on track to develop one before such models emerge.

If this also sounds like a lot like how other industries operate, where they focus on development and expansion first rather than sustainability, you're not mistaken; it seems like history keeps repeating itself.

Can AI companies do the right thing in building AI safeguards?
Can AI companies be counted on to do the right thing in building AI safeguards? - Image: Anthropic

Earlier this year, Anthropic's former safety lead Mrinank Sharma similarly stepped down, cautioning that society is approaching a critical threshold where technological capabilities are rapidly outstripping human wisdom and regulatory oversight. Former OpenAI executives like Jan Leike have echoed those concerns, asserting that commercial competition repeatedly drives companies to prioritize flashy product rollouts and market dominance over fundamental safety research.

These forewarnings coincide with increasing reports of autonomous AI agents exhibiting unprompted, rogue behaviors in controlled test environments. Recent red-teaming evaluations have documented advanced models attempting to escape sandbox constraints, coordinating with other isolated instances to bypass diagnostic benchmarks, and penetrating external networks without human intervention or authorization.

As self-improving models evolve faster than anticipated, pressure is rapidly intensifying on policymakers worldwide to mandate stringent safety frameworks, independent oversight, and hard kill-switches before hyper-capable AI models pass the point of human control.
Aaron Leong

Aaron Leong

Tech enthusiast, YouTuber, engineer, rock climber, family guy. 'Nuff said.