06 · Neural nets · 5 min read · Interactive · updated
How does backpropagation work in neural networks?
In short
Backpropagation computes how the loss depends on every weight in a network by sending an error signal from the output back to the input via the chain rule.
What it is
Backpropagation (backprop for short) is the algorithm for computing the gradient of the loss function with respect to all the weights of a neural network. It applies the chain rule layer by layer, from the output back to the input, and reuses intermediate results stored during the forward pass. As a result, the gradient for millions of weights costs only a few times more than a single forward pass.
Backpropagation on its own does not "learn" anything. It answers only one question: how much would the loss change if this particular weight were nudged up slightly? Learning is the optimizer's job — e.g. gradient descent, which uses that answer to move the weights towards a lower loss.
Intuition: after a lost match, a coach does not change everything at once. They work out who contributed to the result and how much: first the final pass, then the play that led to it, and so on backwards. Backpropagation does the same with the error, dividing "responsibility" among the neurons in proportion to their influence on the outcome.
Mechanism — why it works this way
A network is a composition of functions: the loss L depends on the output y, y on the sum z of the last layer, z on the activation h of the hidden layer, and h on the weights W₁. The chain rule says that derivatives of a composition multiply: ∂L/∂W₁ = ∂L/∂y · ∂y/∂z · ∂z/∂h · ∂h/∂W₁. Backpropagation evaluates this product from the left, i.e. starting from the loss, and reuses each partial result for all the weights of the layer below.
That reuse is the whole secret of its efficiency. The naive method — perturbing each weight slightly and running the forward pass again — needs as many passes as there are weights. With a million weights, that is a million passes. Backpropagation needs one forward pass and one backward pass, because the error signal δ = ∂L/∂z for a given layer is computed once and distributed to all its weights: ∂L/∂W = δ · (layer input).
Passing the signal backwards through a layer takes two operations. Multiplying by the transposed weight matrix distributes the error among the neurons of the previous layer (δ_h = Wᵀ·δ). Multiplying by the activation's derivative f'(z) silences the neurons that were insensitive during the forward pass. For the sigmoid f'(z) = σ(z)(1 − σ(z)) ≤ 0.25, so in a deep network the signal can weaken from layer to layer. This is the source of the vanishing gradient. For ReLU the derivative is either 1 or 0, which eases the problem but creates the risk of "dead" neurons.
Backpropagation is a special case of reverse-mode automatic differentiation. Modern libraries do not require you to derive formulas by hand: they record the graph of operations during the forward pass and traverse it backwards automatically. The idea was discovered independently several times; in neural network training it was popularized by the 1986 paper of Rumelhart, Hinton and Williams.
A limitation: the gradient only tells you how to change the weights so the loss decreases locally, in an infinitesimally small neighbourhood. It guarantees neither a global minimum nor good generalization. That is the job of the optimizer, the initialization and regularization.
By example
A 2-2-1 network with sigmoids and cross-entropy, the same one as in the article on the forward pass: x = (1, 0.5), hidden weights (0.5, −0.3) and (0.2, 0.8), output weights (1, −1), target t = 1. The forward pass gave h = (0.587, 0.646), y = 0.485 and a loss L = 0.723. For a sigmoid with cross-entropy the error signal at the output is remarkably simple: δ = y − t = −0.515. The gradient of the output weights is δ·h = (−0.302, −0.332). The negative sign means that increasing these weights will reduce the loss.
The error travels back to the hidden layer: δ·w = (−0.515, +0.515), multiplied by the sigmoid derivatives h(1 − h) = (0.242, 0.229), gives δ_h = (−0.125, 0.118). The gradient of the weight from x₁ to the first neuron is −0.125·1 = −0.125. A finite-difference check (perturbing the weight by 10⁻⁶) gives −0.1248268, the same value to seven decimal places. One gradient descent step with a learning rate of 0.5 raises the output from 0.485 to 0.613 and lowers the loss from 0.723 to 0.490.
In practice
- PyTorch:
loss.backward()runs backpropagation and adds the gradients toparam.grad; calloptimizer.zero_grad()before every step, because gradients accumulate. - Gradient checking:
torch.autograd.gradcheckcompares the analytical gradient with a finite difference. Always verify your own implementation this way, preferably infloat64. - When gradients explode (the loss becomes
nan), use clipping:torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0). - For monitoring, print the gradient norm of each layer; norms of the order of 10⁻⁸ in the first layers signal a vanishing gradient.
- scikit-learn does all of this inside
MLPClassifier.fit; you get access to the gradients only in libraries such as PyTorch or JAX.
Frequently asked questions
- Is backpropagation the same as gradient descent?
- No. Backpropagation computes the gradient, and gradient descent (or Adam, or SGD with momentum) uses it to change the weights. You can compute a gradient with backpropagation and use it in an entirely different optimization algorithm.
- Does the brain learn by backpropagation?
- Nobody knows, and the literal version is unlikely, because it would require sending the error back along exactly the same connections in the opposite direction. Neuroscientists study biologically plausible approximations, but it remains an open question.
- Why do activations from the forward pass have to be stored?
- The gradient formulas contain values from the forward pass: the layer's input (for the weight gradient) and the activation's derivative at that point. That is why training uses much more memory than prediction. Gradient checkpointing saves memory at the cost of recomputing some of the activations.
Sources
- Rumelhart D. E., Hinton G. E., Williams R. J., "Learning representations by back-propagating errors", Nature 323, 1986, pp. 533–536.
- Goodfellow I., Bengio Y., Courville A., "Deep Learning", MIT Press, 2016, section 6.5 "Back-Propagation and Other Differentiation Algorithms".
- Bishop C. M., "Pattern Recognition and Machine Learning", Springer, 2006, section 5.3 "Error Backpropagation".
- Baydin A. G., Pearlmutter B. A., Radul A. A., Siskind J. M., "Automatic differentiation in machine learning: a survey", Journal of Machine Learning Research 18(153), 2018, pp. 1–43.
- PyTorch documentation, "A Gentle Introduction to torch.autograd": https://pytorch.org/tutorials/beginner/blitz/autograd_tutorial.html