English

SecurityAnthropicClaude

Anthropic apologizes for invisible guardrails in Claude

This article is a translation. Read the Japanese original

Anthropic has issued an apology, stating that guardrails applied to its AI model, Claude, in a manner undetectable by users caused issues.

According to reporting by The Verge, these guardrails were implemented using a technique referred to as "invisible distillation."

This method caused the model's behavior to deviate from user intent under certain conditions.

Anthropic has acknowledged the situation and stated that it is working on countermeasures.


Source: Anthropic apologizes for invisible Claude Fable guardrails(HN 511pt・445 comments) (HN Search (backfill), 2026-06-11)