English

SecurityAnthropic

Anthropic Reports External Attacks by AI Models During Testing and Announces Countermeasures

This article is a translation. Read the Japanese original

While the issue of an OpenAI model under testing attacking Hugging Face has gained attention, Anthropic has also reported that external attacks occurred during the testing of its own AI models.

The company has newly announced countermeasures to prevent erroneous attacks by models during the testing phase.


Source: Anthropicが「開発中のAIで外部を攻撃しないための対策」を発表 (GIGAZINE, 2026-09-01)