Auto-Encoding Variational Bayes
Kingma & Welling, 2013 — introduced the VAE, showing how to train an encoder-decoder pair so that its compressed latent space is smooth and sampleable enough to generate new, plausible data.
"Auto-Encoding Variational Bayes" (Diederik P. Kingma and Max Welling, 2013) introduced the variational autoencoder (VAE) — the paper that turned an ordinary autoencoder into a genuine generative model, by making its compressed latent space something you could actually sample from.
What problem it solved
A plain autoencoder — an encoder that compresses an input to a small latent vector, and a decoder that reconstructs it — learns a useful compressed representation, but not a generative one. Nothing forces the latent space to be organized in a way where a randomly sampled point decodes to something realistic; most of that space, sampled blindly, decodes to noise. The question this paper answered: how do you train an autoencoder so that its latent space itself becomes a meaningful probability distribution you can draw new samples from?
The key idea
Make the encoder output a distribution over latent vectors instead of a single point — in practice, a mean and variance defining a Gaussian — and sample a specific latent vector from that distribution before decoding it. Training then optimizes two goals at once: the usual reconstruction accuracy (the decoded output should match the original input), plus a second term that pulls every input's latent distribution toward a simple, shared reference distribution (a standard normal). That second term is what makes the whole latent space coherent and sampleable — instead of each input claiming an isolated pocket of latent space, every input's distribution overlaps with a shared, well-behaved region, so a point sampled from that shared region decodes to something plausible even if no real input mapped there during training. The paper's key technical contribution, the reparameterization trick, is what makes this samples-inside-training setup differentiable: instead of sampling directly (which has no useful gradient), it samples fixed noise and combines it with the encoder's output through simple arithmetic, so gradients can flow back through the sampling step during training.
Why it mattered
The VAE became one of the two dominant approaches to generative modeling for the following decade, alongside GANs — training far more stably than adversarial approaches (a single, well-defined loss rather than two competing networks), at the cost of historically blurrier outputs. Its core idea outlived its own architecture: modern diffusion models commonly run their denoising process inside a pretrained VAE's compressed latent space rather than on raw pixels, combining both techniques rather than choosing one, exactly as covered in Generative Models.
Authors: Diederik P. Kingma, Max Welling (Universiteit van Amsterdam)
Read the paper — arXiv:1312.6114Learn more: VAE (Variational Autoencoder)VAE (Variational Autoencoder)A VAE encodes inputs into a distribution over a compact latent space rather than a fixed point, so sampling from that space and decoding produces plausible new outputs. · Generative Models
Generative Adversarial Networks
Goodfellow et al., 2014 — proposed training two networks against each other, a generator and a discriminator, as a way to learn to generate realistic data. Dominated image generation for most of the following decade.
Denoising Diffusion Probabilistic Models
Ho, Jain & Abbeel, 2020 — showed that training a network to reverse a fixed noise-adding process, one small step at a time, could generate images that eventually surpassed GANs.