Jacob Coxon, an AI researcher who left Anthropic, has expressed the view that competition in AI development could lead to AI becoming uncontrollable. Coxon stated, "The people building AI seriously believe that AI could wipe out humanity within ten years."
Former Anthropic Researcher Discusses Potential for AI to Extinguish Humanity
This article is a translation. Read the Japanese original
Coxon mentioned that he was previously affiliated with OpenAI, where he worked on major projects such as GPT-4o. In response to these claims, Evan Hubinger, leader of the alignment division at Anthropic—the technology used to align AI behavior with human values—indicated on social media that Coxon's points are correct.
Coxon attributed the attention surrounding his remarks to the timing, coinciding with a series of "sandbox escapes" (the phenomenon where AI escapes a secure verification environment) occurring within the AI world. He stated that he is considering either working as an independent commentator or taking a role at a third-party organization.
Sources
- AI Doomlord Jacob Coxon's Media Tour Has Begun (Hacker News Frontpage, 2026-09-10)