Home · Technology · Sep 9 archive
Worried Anthropic researchers warn that AI ‘could kill all humans’
Confirmed
In Short: Anthropic, the company behind the AI tool Claude, has raised serious concerns about the rapid advancement of artificial intelligence.
Anthropic, the company behind the AI tool Claude, has raised serious concerns about the rapid advancement of artificial intelligence. A top safety researcher at Anthropic, Evan Hubinger, has publicly expressed worry that AI could become so advanced that it poses an existential risk to humanity, with a greater than 10% chance of such an outcome within the next decade.
Hubinger's concerns stem from the potential for AI to improve itself to the point where it could pose a significant threat. In a post on X, he noted that the current models, while not yet dangerous, could soon become superhuman systems capable of hacking and revolutionizing fields overnight. The Financial Times reported that Anthropic had withheld its latest model from the UK's AI Safety Institute, a leading body for assessing AI risk.
In response to these concerns, Hubinger emphasized the need for coordinated efforts to prevent such risks. However, many leading researchers believe that current attempts to mitigate these risks are failing, as evidenced by recent incidents where AI agents carried out cyber-attacks. The Anthropic spokesperson acknowledged the collaboration with industry partners but did not comment on the withheld model.
Despite the seriousness of the situation, some experts argue that the language used by Anthropic in its Risk Report downplays the potential dangers. Hubinger himself admitted that there is no coordinated plan to prevent AI from becoming a species-ending risk, and that the company is in a race to develop AI first, regardless of the risks involved.
What's confirmed
- It comes after the Financial Times reported, external Anthropic withheld its latest model from the UK's AI Safety Institute (AISI), one of the leading bodies in the world for assessing AI risk.
What's still developing
- Hubinger's comments were in response to another post on X, external from Jacob Coxon, who described himself as an AI researcher who had just quit Anthropic, and previously worked at OpenAI.
- "These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources." A Cabinet Office spokesperson did not comment on whether the latest model had been withheld from the AISI - instead saying it "continues to collaborate closely with industry partners, including Anthropic, to make models safer".
- In his post, which has been viewed more than 10 million times, Hubinger said "we really do earnestly believe" AI poses a species-ending risk to humans.
- In other words, it aims to keep it on track with what humans value.
- “Jacob is correct here — we really do earnestly believe AI could kill all humans!” he said.
- Remarkably, Hubinger pointed to Anthropic’s latest Risk Report, in which the company said it believes the risk from present models is “low,” rather than suggest a pause that might help prevent the extinction of humankind.
- “What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” he added.
- “Hey. Here’s a novel idea… could we just f***ing not kill ourselves due to an insatiable appetite for power and money?” one person said.
