Jacob Coxson, a researcher responsible for training AI systems at Anthropic, announced on 2026-09-09 that he had resigned due to the company's insufficient safety measures. Coxson stated that both OpenAI and Anthropic are rapidly advancing the development of self-improving AI—AI capable of autonomously enhancing its own capabilities—without taking responsible action.
In response to this post, Evan Hubinger, who leads the AI safety team at Anthropic, expressed agreement with Coxson's views. Hubinger voiced concern that superintelligence resulting from recursive self-improvement is progressing faster than expected, stating that he personally believes there is a more than 10% probability that AI will cause the death of all humanity within the next 10 years. On the other hand, regarding the risk of current models leading to human catastrophe, he indicated a perception that it is "low," based on the company's risk report.
Source:
- Anthropicの上級安全研究者が「AIが21世紀末までに人類を滅ぼす可能性が10%以上ある」と発言、「Anthropicの安全対策が甘い」と退職する社員も (GIGAZINE, 2026-09-10)