English

News

Reflections on Whether LLM Self-Referentiality is a Byproduct of Training Rather Than Design

This article is a translation. Read the Japanese original

LLMs, such as GPT 5.6 Pro and Fable, are capable of discussing themselves and Gödel's incompleteness theorems more fluently than humans.

However, there is no mechanism for self-referentiality built into any part of their technology stack.

It is argued that no specific design has been implemented within the Transformer architecture, GPUs, or the training process to achieve self-referentiality.

Instead, these abilities have emerged as a byproduct of training the models to speak on a vast range of topics, such as Pokémon and geology.

Douglas Hofstadter once proposed the theory that self-referential "strange loops" are essential to the essence of intelligence.

However, current LLMs achieve conversational intelligence without explicitly incorporating such mechanisms.

This is similar to the process by which axiomatic systems and algorithms in mathematics achieve universality.

It is suggested that self-referential capabilities are naturally derived results of training to be able to discuss any given phenomenon.