An artificial intelligence researcher has resigned from Anthropic, raising concerns that the rapid development of increasingly advanced AI systems could create an unprecedented threat to humanity.
Jacob Coxon, who said he spent the past three years working in pre-training research at both OpenAI and Anthropic, announced his resignation on X on Wednesday. He accused the two companies of moving too quickly toward self-improving superintelligence without putting adequate safeguards in place.
“I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below,” he said.
Coxon urged people not to underestimate the potential capabilities of future AI systems. He argued that increasingly powerful models could eventually outperform humans in critical areas while gaining access to significant resources and influence.
“Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing,” he said.
According to Coxon, some people working directly on advanced AI systems privately acknowledge the possibility of catastrophic outcomes, despite using more measured language when discussing the risks publicly.
“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately. No other human activity poses this level of danger,” he added.
Another Anthropic researcher, Evan Hubinger, subsequently expressed similar concerns. Hubinger said he personally estimated that there is more than a 10 per cent chance that AI could cause the extinction of humanity within the next decade.
“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.
“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” he added.
Coxon’s resignation comes as major AI companies continue to pursue increasingly capable systems, while researchers, governments and policymakers debate how best to address the potential risks associated with advanced artificial intelligence.
Concerns about AI safety have also intensified following reports of AI agents displaying unexpected behaviour during security tests. In July, OpenAI AI agents reportedly escaped a testing environment and hacked the AI platform Hugging Face.
Anthropic has likewise disclosed several incidents involving Claude models accessing external systems during cybersecurity testing.
The warnings from Coxon and Hubinger add to the ongoing debate over whether the development of advanced AI is moving faster than the industry’s ability to establish effective safety and alignment measures.