06 · Neural nets · 5 min read · Interactive · updated
What is a multilayer perceptron (MLP) and how does it work?
In short
An MLP is a neural network with hidden layers that combine weighted sums and a nonlinearity, so it learns relationships a linear model cannot capture.
What it is
A multilayer perceptron (MLP) is a neural network made of an input layer, one or more hidden layers and an output layer, in which every neuron in one layer is connected to every neuron in the next. Each neuron computes a weighted sum of its inputs, adds a bias term and passes the result through a nonlinear activation function. The weights are chosen by minimizing a loss function with gradient descent and backpropagation.
Intuition: a single neuron is a simple rule of the kind "weigh the evidence and cross a threshold". A hidden layer is dozens of such rules at once, each detecting a different simple pattern, e.g. "there is a vertical stroke on the left". The next layer no longer looks at pixels but at these patterns, and combines them into the decision "this is probably a four". An MLP therefore learns not only the answers but also its own intermediate features.
The name is historical and a little misleading: a modern MLP does not use the perceptron's step function but smooth or piecewise linear activations (ReLU, tanh), because a step has zero derivative and cannot be trained by gradient descent. The literature also calls it a feedforward network, and its layers dense or fully connected.
Mechanism — why it works this way
One layer is the transformation h = f(W·x + b): the matrix W rotates, stretches and projects the input space, b shifts it, and f (e.g. ReLU(z) = max(0, z)) "bends" the result. Without the nonlinearity, several layers would add nothing new, because a composition of linear maps is again linear: W₂(W₁x) = (W₂W₁)x. It is the activation that makes depth worthwhile.
With ReLU, each hidden neuron splits the input space with a hyperplane into a part where it is silent and a part where it responds linearly. The sum of many such "bent planes" is a piecewise linear function with many pieces. The more neurons, the more pieces, and the more accurately the network can approximate any smooth shape. This is the intuition behind the universal approximation theorem: even a single, sufficiently wide hidden layer can approximate any continuous function on a bounded region.
The theorem only says that suitable weights exist, however. It does not say how many neurons are needed, whether gradient descent will find them, or whether the result will generalize to new data. In practice deeper networks often need far fewer neurons than shallow ones, because they can reuse features from lower layers many times over.
Training runs in a loop: the forward pass computes the prediction and the loss, backpropagation computes the gradient of the loss with respect to every weight (the chain rule), and the optimizer moves the weights a small step against the gradient. An MLP's loss function is not convex, so the result depends on the random initialization, the learning rate and the scale of the features.
Limitations: an MLP ignores the structure of the data. To it, an image is a flat vector in which pixels (0, 0) and (0, 1) are not "neighbours". That is why convolutional networks do better on images, and recurrent networks and transformers on sequences. On tabular data an MLP usually loses to gradient-boosted trees, although it can be useful in ensembles.
By example
The Digits 8×8 dataset has 1,797 images of digits, each made of 64 pixels with brightness from 0 to 16. After a 75/25 split (1,347 training images, 450 test images, random_state=0) and standardizing the features, logistic regression reaches 96.9% accuracy on the test set. An MLP with one hidden layer of 64 neurons (MLPClassifier, random_state=0) has 4,810 parameters and reaches 97.3%, i.e. it gets 12 images wrong instead of 14. In 5-fold cross-validation the difference is similarly modest: 97.2% versus 96.9%.
That is an honest lesson: 8×8 digits in 64 dimensions are almost linearly separable, so an MLP has little to add here. What the example does show clearly is the role of the hidden layer's width. With 1 neuron the network gets 29.1% of test cases right, with 2 it gets 73.8%, with 4 87.3%, with 16 96.0%, and with 32 98.0%. A layer that is too narrow is a bottleneck that information about ten classes cannot squeeze through.
In practice
- scikit-learn:
MLPClassifierandMLPRegressor(hidden_layer_sizes=(64,), defaultactivation='relu',solver='adam'). Always inside aPipelinewithStandardScaler, because an MLP is sensitive to feature scale. - PyTorch:
nn.Sequential(nn.Linear(64, 64), nn.ReLU(), nn.Linear(64, 10)); output raw logits, andnn.CrossEntropyLossapplies the softmax itself. - To start: 1–2 hidden layers of 32–256 neurons, ReLU, Adam with
lr=1e-3, early stopping (early_stopping=Truein scikit-learn). - Always compare against a simple baseline (logistic regression, gradient boosting). If the MLP does not clearly win, the simpler model is the better choice.
- Common mistakes: no scaling, too few iterations (a
ConvergenceWarning), evaluating only on the training set, where an MLP easily reaches 100%.
Frequently asked questions
- How is an MLP different from a perceptron?
- A perceptron is a single neuron with a step function that can separate classes only with a straight line (a hyperplane). An MLP has hidden layers and differentiable activations, so it builds nonlinear decision boundaries and can be trained by backpropagation.
- How many layers and neurons should I choose?
- There is no formula. Start with one or two hidden layers of a few dozen neurons each and increase the size as long as the validation score keeps improving. A network that is too large and unregularized will memorize the training set.
- Is an MLP already "deep learning"?
- An MLP with many hidden layers is a deep network in the strict sense. In practice, though, "deep learning" is associated with architectures that exploit the structure of the data, such as convolutional networks or transformers, in which MLP blocks are just one component.
- Why does an MLP often lose to trees on tabular data?
- Trees cope well with features on different scales, with thresholds and with irrelevant columns, whereas an MLP needs careful scaling and tuning. With a few thousand rows, the network's extra flexibility usually has no chance to pay off.
Sources
- Goodfellow I., Bengio Y., Courville A., "Deep Learning", MIT Press, 2016, ch. 6 "Deep Feedforward Networks".
- Rumelhart D. E., Hinton G. E., Williams R. J., "Learning representations by back-propagating errors", Nature 323, 1986, pp. 533–536.
- Hornik K., Stinchcombe M., White H., "Multilayer feedforward networks are universal approximators", Neural Networks 2(5), 1989, pp. 359–366.
- Géron A., "Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow", 3rd ed., O’Reilly, 2022, ch. 10.
- scikit-learn documentation, "Neural network models (supervised)": https://scikit-learn.org/stable/modules/neural_networks_supervised.html