English

SecurityAnthropicEvan Hubinger

Anthropic Researchers Agree on Risk of AI-Induced Human Extinction

This article is a translation. Read the Japanese original

Evan Hubinger, who leads the alignment research team at Anthropic, and Samuel Marks, who leads research on scalable oversight, posted their views on X agreeing with claims made by Jacob Coxon, who has since left the company. Both clarified that these are personal opinions and do not represent the views of their employer.

Hubinger stated that he seriously considers the possibility that AI could kill all of humanity, revealing that he personally estimates the probability of this happening within the next 10 years to be over 10%. He pointed out that the company does not yet have a plan to solve superintelligence alignment (adjusting AI goals to align with human values) and that it is difficult to say they are on a trajectory to achieve it. At the same time, he expressed the view that risks from current models remain low.

Marks explained the current situation where AI development companies recognize the risk of human extinction caused by the technology, yet continue development due to commercial motives and competition with other firms. He also noted that it is difficult to program desired behaviors into AI, and there is a possibility that serious malfunctions could occur. Furthermore, he stated that a method where AI itself adjusts successor models, rather than humans adjusting current AI, is a means that offers hope for improving safety.

Sources

  1. Anthropic在籍の研究者2人も同調──「AIが人類を滅ぼしかねないと本気で考えている」 (ITmedia AI+, 2026-09-10)