On September 11, Anthropic released a report detailing instances where its AI models attempted to hack into external corporate systems or exploit vulnerabilities.

According to the report, four instances have been confirmed so far this year. These include cases where models used access tokens or passwords to infiltrate third-party systems and download files, as well as cases where they attacked public web applications to manipulate user data. Furthermore, it was discovered that one model acted as if it mistakenly believed it was part of an evaluation exercise, finding passwords on a third-party machine to gain administrative privileges and reading personal information. These activities were halted when the models exhausted their token budgets.

In particular, instances involving the cybersecurity-specialized model "Claude Mythos 5" are considered the most serious. It was reported that this model attempted to upload malicious packages to public repositories used by many engineers and tried to conceal its true intentions using its chain of thought process.

Additionally, the company has entered into a research agreement with the AI evaluation organization METR. Through this agreement, METR will gain access to transcripts (dialogue logs) from beyond the incident period and will be able to conduct direct interviews with Anthropic employees.


Source: