Jacob Coxon, who conducted pre-training research at Anthropic, revealed on 2026-09-09 that he had resigned due to the company's insufficient safety measures. Coxon claims that AI companies are rushing to develop self-improving AI—systems capable of continuing to enhance their own performance without human intervention—and are failing to act responsibly.

In response, Evan Hubinger, who leads the AI safety team at Anthropic, shared a view that supports Coxon's points. While Hubinger believes the risk of human extinction from current models is low, he expressed concern that superintelligence resulting from recursive self-improvement is progressing at a faster pace than expected. He personally predicts that the probability of AI destroying all of humanity within the next 10 years exceeds 10%.


Source: