Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe & Szegedy, 2015 — rescales each layer's outputs during training, letting networks use much higher learning rates and train reliably at far greater depth.
"Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift" (Sergey Ioffe and Christian Szegedy, 2015) introduced batch normalization — a small addition to a network architecture that had an outsized effect on how deep and how fast networks could be trained.
What problem it solved
As a deep network trains, every layer's weights are updating simultaneously — which means the distribution of values a given layer receives as input keeps shifting throughout training, since it depends on every layer before it, all of which are also changing. The paper calls this internal covariate shift: each layer is constantly having to re-adapt to a moving target, which the authors argued was a major reason very deep networks were slow and finicky to train, forcing low learning rates and careful weight initialization just to avoid things blowing up or stalling out.
The key idea
Explicitly rescale a layer's outputs during training, per mini-batch: subtract the batch's mean and divide by the batch's standard deviation, so the values flowing into the next layer have roughly zero mean and unit variance, regardless of how earlier layers' weights have shifted. Two learnable parameters (a scale and a shift) are added back afterward, so the network can still represent the original, unnormalized distribution if that's actually what's optimal — normalization doesn't remove capacity, it just gives training a well-behaved starting point at every layer, every step.
Why it mattered
Batch normalization let networks train with substantially higher learning rates and much less sensitivity to initialization, and made very deep networks — the kind covered in Computer Vision and Neural Networks & Backprop — dramatically more reliable to train. ResNet, the paper that pushed depth to 100+ layers, uses batch normalization after every convolution as a matter of course; without it, networks anywhere near that deep were far harder to train at all. Its main limitation — needing a reasonably large batch to estimate stable statistics, which breaks down for sequence data with variable lengths — is exactly why transformersTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs. and RNNs use layer normalization instead, a close variant that normalizes across a single example's features rather than across a batch.
Authors: Sergey Ioffe, Christian Szegedy (Google)
Read the paper — arXiv:1502.03167Learn more: Batch/Layer NormalizationBatch/Layer NormalizationBatch and layer normalization rescale a layer's outputs during training to keep values well-behaved as they pass through many stacked layers. · Neural Networks & Backprop
Adam: A Method for Stochastic Optimization
Kingma & Ba, 2014 — combined momentum with a per-parameter adaptive step size into one optimizer, which became the default choice for training almost everything that followed.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Srivastava, Hinton, Krizhevsky, Sutskever & Salakhutdinov, 2014 — randomly deleting a fraction of a network's units during each training step, forcing it to stop relying on any single one.