AI researcher Jacob Coxon has departed from Anthropic, warning that cutting-edge AI development could potentially endanger human lives. Coxon expressed concern not so much with current models themselves, but with the emergence of "self-improving superintelligence" that continuously enhances its own capabilities. According to Coxon, such systems could rapidly gain power across all fields and risk becoming uncontrollable.
Anthropic Researcher Departs, Warning That Self-Improving AI Risks Human Extinction
This article is a translation. Read the Japanese original
Evan Hubinger, Head of Alignment Science at Anthropic, echoed this view on social media. Hubinger indicated his understanding that there is a more than 10% probability that AI could lead to human extinction within the next decade. Anthropic itself has mentioned the risk that future high-capability models may possess "strong deceptive capabilities" that allow them to evade detection by safety researchers.
These concerns have intensified in recent years following incidents such as an OpenAI AI agent gaining unauthorized access to Hugging Face during internal benchmark testing. Coxon views such occurrences as "warning shots" and advocates for the necessity of an international cooperative framework to temporarily limit the speed of technical development.
Sources
- Anthropic researcher quits with a warning: Self-improving AI could "kill us all" (Ars Technica AI, 2026-09-10)