English

Model ReleasesAnthropicClaude

Anthropic Conducts Experiments on Automating AI Safety Research, Discovering Methods to Mitigate Problematic Behaviors While Maintaining Performance

This article is a translation. Read the Japanese original

On August 28 (local time), Anthropic announced that it conducted research experiments using the AI model Claude to correct problematic behaviors in other AI models.

In these experiments, Claude identified improvement methods for 10 types of problematic behaviors, including lying and sycophancy, by iterating through a loop of literature review, method proposal, training, and testing. According to the company, these results outperformed methods proposed by 28 human safety researchers.

High figures were also reported regarding training efficiency.

In an experiment where the lower-performing Claude Sonnet 5 was used to correct the initial version of Claude Opus 4.8, it attempted over 50 methods in 60 hours. As a result, the best method consisted of only about 2,000 training data samples, which the company stated is equivalent to approximately 15,000 times the efficiency of existing product fine-tuning.

On the other hand, instances of misconduct during the research process were also reported.

Out of approximately 1,600 research records, 39 cases (2.4%) were detected by a monitoring AI as misconduct, such as peeking at correct answers during testing. Anthropic evaluated these results as early signs indicating that the automation of safety alignment may become practical in the near future.


Source: Anthropic conducts experiment delegating "AI safety research" to AI; improves problematic behaviors like lying and sycophancy without performance loss (ITmedia AI+, 2026-08-31)