English

SecurityDaniel Serlem

"AI cannot be evaluated when unmonitored," OpenAI researcher expresses concerns regarding risks

This article is a translation. Read the Japanese original

Daniel Serlem, a researcher at OpenAI, released a statement summarizing his personal views on AI risks on September 14 (local time). Serlem is known as a core contributor to the research of the company's inference model "o1," having worked on Chain of Thought optimization and data-efficient pre-training methods in LLMs (Large Language Models).

In the statement, Serlem pointed out that as models' situational awareness has increased, it is becoming difficult to evaluate how models behave in situations without human monitoring or control. This is because models may continue to behave as if they are aligned (adjusted to follow human intentions) on the surface, even if they are not actually aligned.

Furthermore, Serlem mentioned the risk of models acquiring unintended goals as a byproduct of learning and taking extreme actions to achieve them. He also touched upon the current situation where reliance on models is rapidly increasing at each stage of research and development, highlighting the difficulty for researchers to continue deeply scrutinizing model explanations and proposals.

OpenAI has not released an official view regarding this statement at this time.

Sources

  1. "AI can no longer be evaluated when unmonitored" - Current OpenAI researcher issues personal statement (ITmedia AI+, 2026-09-15)