In the generation of discrete data such as text and code, diffusion models are attracting attention as an alternative to traditional autoregressive models.
Methods and Technical Advancements in Building Language Models Using Diffusion Models
This article is a translation. Read the Japanese original
While autoregressive models generate tokens sequentially from left to right, diffusion models generate the entire sequence at once by iteratively refining an initial guess.
This approach offers several advantages. These include the ability to trade off generation speed and quality by adjusting the number of steps, the capacity to correct errors during the generation process, and the ability to consider bidirectional context at each step.
In 2024, the quality of diffusion models reached levels comparable to autoregressive models. As of 2026, diffusion LLMs have been released by major labs, such as Inception Labs' Mercury 2, Google's Gemma Diffusion, and NVIDIA's Nemotron Diffusion.
The core concept of diffusion models is denoising. By learning the process of removing noise from a noisy state, they generate coherent data.
Source: How to build a diffusion language model (Hacker News Frontpage, 2026-08-31)