English

News

Technical Analysis Released: LLMs Are More Than Just Next-Token Predictors

This article is a translation. Read the Japanese original

A technical analysis has been released arguing that viewing LLMs merely as "next-token predictors" is an incomplete perspective.

During the pre-training stage, models are trained to predict the token that follows an existing sequence contained within the training data. This process focuses on reproducing the next move present in the dataset.

However, post-training performed in modern LLMs—particularly Reinforcement Learning from Verifiable Rewards (RLVR)—is fundamentally different in nature.

In RLVR, the model performs "exploration" by generating new sequences itself and progresses through learning based on the rewards obtained from those results.

The author points out that this process can be compared to a chess engine. While a system that imitates existing game records is a "next-move predictor," a system that selects moves leading to victory through its own exploration cannot be called a mere predictor.

In this way, models that have undergone post-training are said to learn not only by imitating existing text but also from the achievements gained through their own exploration.


Source: Stop Thinking of LLMs as Next-Token Predictors (Hacker News Frontpage, 2026-09-05)