Explaining and Harnessing Adversarial Examples
Goodfellow, Shlens & Szegedy, 2014 — argued that tiny crafted perturbations fool networks because of their near-linear behavior in high dimensions, and gave a one-step method to generate them cheaply.
"Explaining and Harnessing Adversarial Examples" (Ian Goodfellow, Jonathon Shlens, and Christian Szegedy, 2014) is the paper that turned a puzzling observation into a method and an explanation, and launched adversarial robustness as a research area.
What problem it solved
A year earlier, Szegedy and colleagues had reported something strange: adding a tiny, structured perturbation to an image could make a well-trained classifier confidently output the wrong label, even though a person couldn't see any change. That raised an unsettling question nobody could answer. Was this a quirk of a particular network, a byproduct of overfitting, or something deeper? The leading guesses blamed the networks' extreme nonlinearity, which would have made the problem look both exotic and unfixable.
The key idea
The paper's explanation is the opposite of the leading guess: the cause is linearity. In a high-dimensional input, a perturbation that changes each pixel by a tiny, imperceptible amount can still shift the network's output enormously, because the effect of all those small changes adds up across thousands of dimensions. Networks are trained to be roughly linear in many regions precisely because linear behavior is easy to optimize, so they inherit this weakness.
That explanation produced a simple attack, the fast gradient sign method (FGSM): compute the gradient of the loss with respect to the input (not the weights), take the sign of each component, and move the input a small step ε in that direction. One gradient computation, no iteration. It's the same machinery as gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient. turned around, run against the input and in the direction that increases the loss. The paper also showed that training on adversarial examples as they're generated (adversarial training) made networks measurably more robust, a regularizing effect beyond just defense.
Why it mattered
It gave the field both a cheap way to attack and a way to defend, and it established the pattern behind adversarial examplesAdversarial ExampleAn adversarial example is an input with a small, often imperceptible perturbation deliberately crafted to make a model produce a wrong or attacker-chosen output. as the lesson describes them: perturbations aimed at the model's own decision boundary, derived from its own gradients. The paper also noted that adversarial examples often transferred between models trained separately, which is what makes them a practical threat even when an attacker can't see the target's weights. The problem has proven stubborn: robustness defenses keep being broken, and the same gradient-guided idea reappears in attacks on language models, including automatically generated suffixes that push a model past its safety training, covered in AI Security.
Authors: Ian J. Goodfellow, Jonathon Shlens, Christian Szegedy (Google)
Read the paper — arXiv:1412.6572See also: Ian GoodfellowIan GoodfellowInvented generative adversarial networks (GANs) in 2014, the architecture that first made realistic image generation possible by pitting two networks against each other.
Learn more: Adversarial ExampleAdversarial ExampleAn adversarial example is an input with a small, often imperceptible perturbation deliberately crafted to make a model produce a wrong or attacker-chosen output. · AI Security · Generative Adversarial Networks
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Shazeer et al., 2017 — a layer made of thousands of expert subnetworks where a learned router activates only a few per input, growing parameter count enormously without growing compute.
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Bricken et al., 2023 (Anthropic) — trained a sparse autoencoder on a small transformer's activations and recovered thousands of human-interpretable features that individual neurons mix together.