Denoising Diffusion Probabilistic Models
Ho, Jain & Abbeel, 2020 — showed that training a network to reverse a fixed noise-adding process, one small step at a time, could generate images that eventually surpassed GANs.
"Denoising Diffusion Probabilistic Models" (Jonathan Ho, Ajay Jain, and Pieter Abbeel, 2020) is the paper that made diffusion models practical, kicking off the approach that now dominates state-of-the-art image generation.
What problem it solved
Diffusion-style ideas — destroy structure with noise, then learn to reverse it — had been proposed years earlier, but early versions were difficult to train and produced results well behind GANs and VAEs. The open question was whether a diffusion-style model could be made to actually work well in practice, with a training objective simple and stable enough to compete with the dominant generative approaches of the time.
The key idea
Fix the forward "destroy" process as pure, un-learned arithmetic: gradually add small amounts of Gaussian noise to a real image over many steps, until it's indistinguishable from pure noise. Then train a single network to do one narrow, well-defined job at every step of the reverse direction: given a noisy image and how many steps of noise it's had added, predict the noise that was added — not the clean image directly, a subtlety the paper showed mattered for training stability. Subtracting that predicted noise moves the image one step back toward clean. The training objective this produces is strikingly simple: an ordinary mean-squared-error regression between the network's predicted noise and the actual noise that was added, no adversarial setup and no sampling required during training at all. Generating a new image just runs this trained denoising step repeatedly, starting from pure random noise.
Why it mattered
This paper's specific parameterization (predict the noise, not the image) and training recipe are what turned diffusion from a theoretical curiosity into a method that produced genuinely competitive, eventually superior, image quality compared to GANs — with none of GANs' notorious training instability, since there's no adversary and no Nash equilibrium to hunt for, just a single stable regression loss. The tradeoff, covered in Generative Models, is inference cost: generating one image means running the denoising network many times in sequence, far slower than a GAN's single forward pass. Nearly every major image (and increasingly video and audio) generation system since — Stable Diffusion, Midjourney's underlying approach, and others — builds directly on the training recipe this paper established.
Authors: Jonathan Ho, Ajay Jain, Pieter Abbeel (UC Berkeley)
Read the paper — arXiv:2006.11239Learn more: Diffusion ModelDiffusion ModelA diffusion model learns to reverse a fixed process of gradually adding noise to data, generating new samples by denoising pure noise step by step. · Generative Models
Auto-Encoding Variational Bayes
Kingma & Welling, 2013 — introduced the VAE, showing how to train an encoder-decoder pair so that its compressed latent space is smooth and sampleable enough to generate new, plausible data.
Playing Atari with Deep Reinforcement Learning
Mnih et al., 2013 (the "DQN" paper) — combined Q-learning with a neural network to learn to play Atari games directly from raw pixels, no hand-engineered features, from a single reward signal.