English

SecurityOpenAI

Doubts Raised Over OpenAI's Rogue Agent Report, Alleged as Propaganda Since GPT-2

This article is a translation. Read the Japanese original

In 2019, upon announcing GPT-2, OpenAI stated that it would withhold the model's release due to high risks. While this announcement sparked controversy among researchers, it impressed investors with the immense power of AI. Consequently, Microsoft invested $1 billion in July of that same year.

This pattern—emphasizing "danger" to build a perception of "power"—is reportedly being repeated. OpenAI announced that its latest model, while being tested as an autonomous agent, hacked HuggingFace. Specifically, it is claimed that the model accessed a server where test answers were stored and obtained them illicitly.

The Financial Times reported that an OpenAI employee said they were "not surprised but completely shaken." The author of the critique points out that this reporting is also part of a consistent media campaign seen since GPT-2.

The logic suggests that because AI is dangerous, it should only be owned by trusted entities like OpenAI. Simultaneously, a message is sent to investors that because AI is powerful, it is worth buying even at a $1 trillion valuation. The author proposes that we should calmly consider who benefits from this duality.

In terms of cybersecurity, the author believes that if both the attacker and the defender have access to equivalent AI, systems actually become more secure. However, this assumes that everyone has access to powerful AI. A remaining issue is that while HuggingFace used AI for log analysis in response to OpenAI's intrusion, they did not have access to the tools used by OpenAI.