Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Srivastava, Hinton, Krizhevsky, Sutskever & Salakhutdinov, 2014 — randomly deleting a fraction of a network's units during each training step, forcing it to stop relying on any single one.
"Dropout: A Simple Way to Prevent Neural Networks from Overfitting" (Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, 2014) introduced dropout — a strikingly simple regularization technique that became a default ingredient in large networks for years afterward, including AlexNet, published by three of the same authors two years earlier.
What problem it solved
Large neural networks have enough capacity to memorize their training data outright rather than learn patterns that generalize — overfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data.. A specific way this shows up: units in a network can end up in co-adaptations, where several units only work correctly in combination with each other's specific idiosyncrasies, a fragile, overly-specific reliance that doesn't transfer to new data. Existing regularization techniques at the time didn't directly target this failure mode.
The key idea
During training, randomly and independently "drop" each unit (set its output to zero) with some fixed probability, typically 50%, at every training step. A different random subset of the network is deleted on every single step, which means no unit can safely rely on any other specific unit being present — the co-adaptation the paper targeted becomes much harder to form, since the exact set of collaborators a unit can lean on keeps changing underneath it. At test time, dropout is turned off and every unit's output is scaled down to compensate for the fact that, during training, only a fraction of them were active on average. The authors also frame this as training an enormous number of different "thinned" sub-networks that share weights, and test time as approximating the average prediction across all of them — a cheap approximation to model averaging, normally an expensive technique requiring many separately trained models.
Why it mattered
Dropout was cheap to implement, added no meaningful computational overhead, and reliably improved generalization across a wide range of architectures and tasks, making it one of the most widely adopted regularization techniques of the deep learning era — AlexNet used it in its fully-connected layers as one of the specific choices credited with making its result possible. Later architectural changes reduced how load-bearing dropout is in some settings — batch normalization and residual connections both have some regularizing effect of their own, and very large models trained on enormous datasets sometimes skip dropout entirely — but it remains a standard, well-understood tool covered alongside other regularization techniques in ML Fundamentals.
Authors: Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, Ruslan Salakhutdinov (University of Toronto)
Read the paper — JMLR, Vol. 15See also: Geoffrey HintonGeoffrey HintonKnown as a godfather of deep learning — co-authored backpropagation in 1986, co-invented the Boltzmann machine and dropout, and co-authored AlexNet, the 2012 result that restarted the field. · Alex KrizhevskyAlex KrizhevskyBuilt AlexNet in 2012 with Ilya Sutskever and Geoffrey Hinton, the deep convolutional network whose ImageNet win is widely credited with kicking off the deep learning boom. · Ilya SutskeverIlya SutskeverCo-authored AlexNet as a student, then sequence-to-sequence learning, then co-founded OpenAI and helped drive the bet that scale would produce GPT-level language models.
Learn more: RegularizationRegularizationRegularization is any technique that trades some training-data fit for better generalization, fighting overfitting on purpose. · ML Fundamentals
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe & Szegedy, 2015 — rescales each layer's outputs during training, letting networks use much higher learning rates and train reliably at far greater depth.
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, Sutskever & Hinton, 2012 — the paper widely credited with restarting deep learning, by winning ImageNet with a deep CNN trained on GPUs by a margin nobody expected.