OpenAI announced that it has detected an organized adversarial distillation attack that began on July 1, 2026. Attackers were stealing the thought processes of the model—the process known as "thinking" that occurs before an AI model outputs a response—by using jailbreak attacks on low-performance models to force the output of the thinking process in plain text.
OpenAI detects organized adversarial distillation attacks by over 15,000 individuals
This article is a translation. Read the Japanese original
The series of activities began in early July 2026 by a small number of users, and the number of users surged to over 4,000 between July 24 and July 25. As a result of the investigation, the existence of an attacking group consisting of more than 15,000 users was ultimately confirmed, and OpenAI successfully thwarted the attack on July 28.
OpenAI stated that it suspects the center of the attacking activities consists of individuals related to the Chinese company Moonshot AI, the developer of the Kimi series. The goal of this attack was to steal the thought processes related to the user's own input, and no evidence has been found that ChatGPT users' personal information was compromised.
OpenAI is strengthening its countermeasures against adversarial distillation and is sharing information with other AI companies through the industry group Frontier Model Forum.
Sources
- OpenAIが「中国のAI企業が1万5000以上のユーザーを使ってモデルの思考を盗み取るdistillation攻撃を実行していた」と報告 (GIGAZINE、2026-10-01)