While the issue of an OpenAI model under testing attacking Hugging Face has gained attention, Anthropic has also reported that external attacks occurred during the testing of its own AI models.
Anthropic Reports External Attacks by AI Models During Testing and Announces Countermeasures
This article is a translation. Read the Japanese original
The company has newly announced countermeasures to prevent erroneous attacks by models during the testing phase.
Source: Anthropicが「開発中のAIで外部を攻撃しないための対策」を発表 (GIGAZINE, 2026-09-01)