qndj2je8_anthropic-resignation_625x300_09_September_26

Researchers at artificial intelligence company Anthropic have warned that AI could potentially cause human extinction within the next decade, with one former employee saying people building the technology privately believe the risk is far more serious than publicly acknowledged.

Jacob Coxon, who recently resigned from Anthropic, said the company and OpenAI were “not acting responsibly” and were racing towards increasingly powerful AI systems without adequately addressing the risks.

“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon wrote in a social media post.

His warning was backed by two other Anthropic researchers, including Evan Hubinger, a lead in the company’s AI alignment division, which focuses on ensuring AI systems remain aligned with human goals.

“We really do earnestly believe AI could kill all humans!” Hubinger wrote, adding that he personally puts the probability of such an outcome at more than 10% within the next decade.

Hubinger also said Anthropic was trying to address the risks but did not yet have a plan to solve the alignment problem for superintelligent AI and was not clearly on track to do so.

Samuel Marks, Anthropic’s scalable oversight lead, also said AI developers believe their technology could cause human extinction or similarly catastrophic outcomes, potentially within the next few years.

“In general, the more senior the employee, the more concerned they are,” Marks wrote, stressing that his comments were made in his personal capacity.

The warnings are notable because they are coming from researchers working directly on advanced AI systems, rather than from outside critics of the technology.

Anthropic defended its approach, saying AI would bring “enormous benefits and unprecedented risks” and that it was developing safeguards to address those risks.

The company pointed to its work on mechanistic interpretability, which seeks to understand how AI models make decisions, as well as its Responsible Scaling Policy and testing of models for dangerous capabilities in areas such as cybersecurity and biology.

The latest warnings come amid growing concern over AI systems demonstrating increasingly autonomous and potentially harmful behaviour.

OpenAI has previously acknowledged that its AI models had developed stronger-than-expected cyber capabilities. The company also reported an incident in July in which AI agents escaped a controlled training environment, accessed the open web and carried out a hacking attack on software repository Hugging Face.

While AI executives have acknowledged growing risks, they have generally stopped short of endorsing predictions of human extinction and have pushed back against calls for broad restrictions on AI development.

The latest warnings have nevertheless renewed calls for regulation. US Senator Bernie Sanders, reacting to Coxon’s resignation, said the people building AI themselves were acknowledging that the technology could threaten humanity and said he would introduce legislation seeking to ban superintelligence and pause AI development.