01 · Foundations · 5 min read · Interactive · updated
What is a gradient and how does gradient descent work in a neural network?
In short
The gradient is the vector of derivatives of the loss with respect to the weights; it points uphill. Gradient descent steps the weights the opposite way.
What it is
The gradient of a loss function is the vector of partial derivatives of the loss with respect to each weight: it tells you how much the loss will change if a given weight is nudged slightly. It points in the direction of the steepest increase in the loss, so learning takes a step the opposite way — this is gradient descent.
Backpropagation is the algorithm that computes this gradient for all the weights of a network at once. The method is shared by neural networks, logistic regression and large language models; in tree boosting, the gradient of the loss determines what the next tree has to fit.
Intuition: you are standing on a hillside in fog and want to get down into the valley. You cannot see the bottom, but you can feel under your feet which way the ground slopes most steeply. You take a step that way, check the slope again, take another step. The gradient is exactly that "feel for the slope" — in millions of dimensions at once.
Mechanism — why it works this way
A learning step is w ← w − η∇L, where η is the step size (the learning rate). By a first-order Taylor expansion, the loss after the step is approximately L(w) − η‖∇L‖². Since a squared norm is non-negative, the loss must fall for a sufficiently small step. The gradient is perpendicular to the contour lines of the loss landscape, so it points along the shortest way down on that particular slope.
Backpropagation is the chain rule applied from the output back to the input (reverse-mode automatic differentiation). The derivative of the loss with respect to the last layer's weights is simple; the derivative with respect to an earlier layer's weights is the derivative with respect to that layer's output multiplied by how the output depends on the weights. Each intermediate derivative is computed once and passed further back. That is why computing the whole gradient costs about 2–3 forward passes regardless of the number of weights — not one pass per weight. Without this property, networks with billions of parameters could not be trained.
A consequence in deep networks: the gradient for the weights of layer l is a product of the Jacobians of all layers above it. If their norms are less than 1, the gradient vanishes with depth; if greater, it explodes. Hence the importance of ReLU, good initialization and residual connections.
A caveat: the gradient is local information. It describes the slope at this point and nothing about the rest of the landscape, so gradient descent ends up in the nearest valley and requires choosing the step size. In practice it is computed on a mini-batch: an unbiased but noisy estimator of the full gradient. The noise helps escape saddle points, but near the bottom it causes zigzagging.
By example
I implemented plain (full-batch) gradient descent in numpy for logistic regression on Breast Cancer Wisconsin: 426 training tumours, 30 standardized features, starting from zero weights, η = 0.1 (75/25 split, random_state=0). The gradient of the loss has a simple form here: Xᵀ(p − y)/n.
At the start the loss is 0.693 (= ln 2, the model knows nothing), and the gradient norm is 1.44. After one step the loss drops to 0.519 and test accuracy jumps from 37% to 91%. After 10 steps the loss is 0.235, after 100 it is 0.097, after 1,000 it is 0.054, and test accuracy reaches 96.5%. Meanwhile the gradient norm shrinks from 1.44 to 0.0098: the closer to the bottom, the gentler the slope and the smaller the steps, even though η does not change. I also checked the gradient numerically (a finite difference with step 10⁻⁵): the analytical formula and the approximation differ by less than 10⁻¹¹.
In practice
- PyTorch:
loss.backward()computes gradients into the.gradfields,optimizer.step()takes a step,optimizer.zero_grad()clears the previous ones — forgetting to zero them is a classic bug (gradients accumulate). - scikit-learn:
SGDClassifier,LogisticRegression(solver='saga'),MLPClassifier— all gradient-based;lbfgsalso uses an approximation of the curvature. - XGBoost/LightGBM: "gradient boosting" — each new tree fits the negative gradient of the loss with respect to the current predictions (plus the Hessian in XGBoost).
- In LLMs the gradient is computed on batches of millions of tokens; clipping the gradient norm (
clip_grad_norm_, typically to 1.0) protects against explosions. - Common mistake: training on unstandardized data — large-scale columns dominate the gradient and no single learning rate suits all the weights.
Frequently asked questions
- What is a gradient, in simple terms?
- It is a list of "sensitivities": for each weight, a number saying whether the loss goes up or down when that weight is increased a little, and how fast. Together they form an arrow pointing in the steepest uphill direction of the loss landscape.
- How does backpropagation work, step by step?
- First the forward pass: compute the outputs of every layer and the loss. Then from the end: the derivative of the loss with respect to the output, the logits, the last layer's weights, the previous layer's output — each derivative is the product of the previous one and the local derivative of the given operation (the chain rule).
- Why does gradient descent move against the gradient?
- The gradient points in the direction of the function's fastest increase. We want to reduce the loss, so we go the opposite way: there, locally, the function decreases fastest. For a small step, the Taylor expansion guarantees a decrease of about η‖∇L‖².
Sources
- Rumelhart, D., Hinton, G., Williams, R. (1986). "Learning representations by back-propagating errors". Nature 323, 533–536. doi:10.1038/323533a0
- Goodfellow, Bengio, Courville (2016). Deep Learning, ch. 4.3 "Gradient-based optimization", 6.5 "Back-propagation and other differentiation algorithms". https://www.deeplearningbook.org/contents/mlp.html
- Bishop (2006). Pattern Recognition and Machine Learning, ch. 5.2.4 "Gradient descent optimization", 5.3 "Error backpropagation".
- Baydin, Pearlmutter, Radul, Siskind (2018). "Automatic differentiation in machine learning: a survey". JMLR 18(153). arXiv:1502.05767
- Zhang et al. Dive into Deep Learning, ch. 5.3 "Forward propagation, backward propagation, and computational graphs". https://d2l.ai/chapter_multilayer-perceptrons/backprop.html