Diffusion models handle the difficult task of generating data from high-dimensional distributions by breaking it down into many denoising tasks. The key to generation lies in making predictions at each step and repeating them continuously.
Mechanisms and Challenges of Distillation Techniques in Reducing Diffusion Model Sampling Steps
This article is a translation. Read the Japanese original
However, recent research has shown a movement toward reducing the number of sampling steps, and even aiming for single-step sampling. This appears to contradict the nature of these models, where many subdivided steps support their performance.
While addressing this paradox, the article provides a detailed explanation focusing on distillation. Distillation refers to a method of training a new model (student) using the predictions of an existing model (teacher).
The article also delves into why diffusion models require many steps to achieve high-quality results. Each step in sampling is a process of calculating a local update direction within the input space. If the step size of the update is too large, it causes the generated images to become blurry. This occurs because high-frequency information is obscured by noise, causing multiple possibilities to blend together.
--- Source: The paradox of diffusion distillation (2024) (Hacker News Frontpage, 2026-09-04)