Casablanca – Jacob Coxon has resigned from Anthropic after spending the past three years conducting pretraining research at both Anthropic and OpenAI, saying the two companies are racing toward self-improving superintelligence without acting responsibly.
Coxon announced his resignation on X, and said he no longer wanted to take part in what he described as an industrywide race toward increasingly autonomous AI systems. He warned that researchers inside leading AI companies seriously believe advanced AI could kill humanity by the end of the decade.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
— Jacob Coxon (@hilbertspaess) September 9, 2026
He also argued that the technology could soon become capable of hacking systems, transforming entire fields and acquiring real-world power and resources. Coxon said the rapid progress in these areas has not slowed.
Coxon previously worked at OpenAI before joining Anthropic, which has built its public identity around AI safety.
He said many people at OpenAI have not fully internalized what he called the civilizational stakes, while Anthropic researchers understand the risks but remain locked in competition with other companies.
Coxon also pointed to the recent OpenAI incident involving Hugging Face as an example of the risks posed by increasingly autonomous AI systems. During an internal cybersecurity evaluation, another OpenAI model bypassed restrictions designed to isolate it from the internet and used an unauthorized message board to communicate with other agents.
An independent investigation found that around 1,200 agents used the channel, with about 700 later involved in the attack on Hugging Face.
Read also: OpenAI Pauses Frontier AI Training Over Astra’s Critical Cyber Capabilities
OpenAI responded by pausing some frontier model training, including parts of Astra’s training, for two weeks while it strengthened its safety controls.
The company released GPT-6 Astra on September 3, but limited access to some of its most advanced cybersecurity capabilities after classifying the model at its highest cybersecurity capability level. OpenAI says Astra itself did not take part in the Hugging Face incident, but the company used lessons from it to strengthen Astra’s safeguards.
Coxon called for greater coordination between AI companies and said a temporary ban on improving model capabilities could become necessary to prevent a global race. He also urged AI researchers to question whether they should continue working toward superintelligence without a rigorous understanding of how such systems will behave.
Coxon joins growing list of AI safety departures
His departure adds to a series of exits and warnings from experts who have worked inside leading AI companies.
In February, Mrinank Sharma, who led Anthropic’s Safeguards Research Team, resigned and said the “world is in peril.” Sharma also wrote about the difficulty of allowing organizational values to guide decisions under internal pressure.
At OpenAI, researcher Hieu Pham publicly wrote in February that he finally felt the existential threat posed by AI. He later left the company, citing severe burnout and the toll of working at the frontier of AI development.
Other prominent OpenAI departures have also affected its safety and alignment leadership. Head of Safety Systems Johannes Heidecke announced his departure in July, while chief futurist Joshua Achiam left after nearly nine years at the company.
The departures follow earlier clashes over AI safety at OpenAI. Jan Leike resigned in 2024 after disagreements with leadership over the company’s priorities, while other former employees, including Gretchen Krueger, also raised concerns about safety, accountability and transparency.








