Jacob Coxon is an artificial intelligence researcher who had worked training AI models at OpenAI and has just resigned from Anthropic, concerned about the dangerous way the technology is developing. “I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are rushing straight towards a self-improving superintelligence and betting with our lives,” he posted on his X account.
Read more Civil Protection lifts the mobility restriction in the five regions affected by the storm
The engineer issued a warning: “do not underestimate the power of this technology. Soon they will be superhuman systems capable of hacking anything, revolutionizing any field overnight, and acquiring real power and resources.” According to Coxon, “the people building AI sincerely believe it could kill us all by the end of the decade.” “This is not a marketing trick — he warned —. If anything, many executives and senior researchers moderate their language in the press to sound sensible, but I hear the same people express fear in private. No other human activity represents this level of danger.”
“Accepting this race and entering the ‘endgame’ — final game — is an arrogant bet that should not be made from the Slack of a private company”
Although Coxon has worked at both companies, he considers Anthropic to have a more responsible attitude. “At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well understood, but they are trapped in a race to be first: they believe no one else will act responsibly, so they must do it themselves, despite the risk,” he said. In his opinion, “accepting this race and entering the endgame — final game — is an arrogant bet that should not be made from the Slack of a private company.”
The resigning Anthropic employee considers himself “optimistic about the potential for coordination” because warning signs like the attack on Hugging Face that an OpenAI AI carried out on its own “have made pace agreements between U.S. labs more viable.” Despite this, Coxon does not consider “that we are on track to prevent a global race, which could require costly actions such as a temporary ban on improving model capabilities.”
Read more The two Catalonias
His final reflection on how he made the decision to resign and explain his concerns to the world points out: “if you are a lab researcher, I urge you to consider how the coming years will really feel. Do you want to start a superintelligent reinforcement learning run without a rigorous understanding of its mind? Should you lower your head because ‘anyway, it’s happening’ or take this moment to ask for different conditions?”
One of Coxon’s colleagues, Evan Hubinger, who researches stress testing of AI models at Anthropic, has reaffirmed his concerns. “Jacob is right here we really sincerely believe that AI could kill all humans.”
Also read
“Personally — added Hubinger —, I think it is >10% within the next decade.” “I think Anthropic is doing the best it can, but we still don’t have a plan to solve alignment for superintelligence and we are not clearly on track to achieve it,” he said. According to this engineer, the most worrying thing is that AI constantly improves itself: “what worries me is the superintelligence that arises from recursive self-improvement, as we have said it is happening faster than we thought.”