04 · Evaluation · 5 min read · Interactive · updated
What is overfitting and how do you spot it?
In short
An overfit model memorizes noise in the training set and does worse on new data; an underfit model is too simple to capture even the main pattern.
What it is
Overfitting is when a model has fitted the training set so closely that it has also learned its accidental features — noise, exceptions, label errors — and as a result predicts worse on new data. Underfitting is the opposite: the model is too simple, or trained too briefly, to capture even the main relationship, so it is wrong both on the training data and on new data.
Intuition: a student who memorized the answers to last year's exam will score 100% on it but fail a new set of questions — that is overfitting. A student who only remembered "the answer is usually B" will do poorly on both exams — that is underfitting. We want someone in between: a student who understood the principles and can therefore handle questions they have never seen.
The key point is that overfitting cannot be seen on the training data. There, the overfit model looks best. It reveals itself only when you compare the training score with the score on data the model has not seen.
Mechanism — why it works this way
Every training set is only a sample. It contains the real signal (e.g. "larger tumours are more often malignant") and the randomness of that particular sample (that one patient with unusual values). A learning algorithm minimizes training error and has no way to tell signal from noise — to it, both are simply patterns to fit.
A model with high capacity (a deep tree, a network with millions of weights, a high-degree polynomial) can fit almost any arrangement of points, so it fits the noise too. But the noise in new data is different, so those "learned" details stop fitting and add error. A low-capacity model (a straight line through clearly curved data, a tree with a single split) cannot accommodate even the signal — that is underfitting.
In statistical terms this is the bias–variance trade-off. An underfit model has high bias: it is systematically wrong in the same way. An overfit model has high variance: trained on a different sample from the same population, it would give noticeably different predictions. The error on new data is roughly the sum of these two components plus irreducible noise, which is why the minimum lies somewhere in the middle of the complexity scale.
The risk of overfitting grows when there is little data relative to the number of parameters, when there are many features (making it easier to find a spurious correlation), when labels are noisy and when training runs too long. It shrinks with more data, regularization, a simpler model, early stopping and averaging many models.
A caveat: the picture of a "U-shaped" error curve is not always complete. In very large networks one observes double descent — past the threshold at which the model exactly interpolates the data, test error can start falling again. This does not invalidate the notion of overfitting, but it shows that "number of parameters" is an imperfect measure of complexity.
By example
I split the Breast Cancer Wisconsin dataset (569 tumours, 30 features, positive class = malignant) with stratification into 426 training and 143 test examples (train_test_split, test_size=0.25, random_state=0) and trained decision trees of increasing maximum depth. A tree of depth 1 (one split, 2 leaves) has 92.5% accuracy on the training set and 87.4% on the test set — that is underfitting: it makes quite a few mistakes in both. A tree of depth 3 (8 leaves): 97.4% on training and 95.8% on test.
A tree with no depth limit grows until every leaf is pure (18 leaves) and reaches 100% on training, but only 94.4% on test — less than the three-level tree. Those extra leaves describe individual unusual cases from the training sample, not a rule. The gap between the training and test scores (here 5.6 percentage points) is the basic signal of overfitting. On a test set of 143 examples, one error is worth 0.7 points, so small differences between depths 3–5 fall within the noise — which is why depth is chosen by cross-validation rather than a single split.
In practice
- Always compare the training score with the validation score:
cross_validate(model, X, y, return_train_score=True). High training and low validation = overfitting; both low = underfitting. - Plot a validation curve (
ValidationCurveDisplay) for a complexity parameter such asmax_depth,Coralpha, and a learning curve (LearningCurveDisplay) to check whether more data would help. - Remedies for overfitting: regularization (
Ridge,CinLogisticRegression), constraining the tree (max_depth,min_samples_leaf), dropout and weight decay in PyTorch, early stopping, more data, augmentation. - Remedies for underfitting: a richer model, extra features (e.g.
PolynomialFeatures), weaker regularization, longer training. - Common mistake: picking the model "to suit the test set" through repeated attempts — that is overfitting too, just at the level of the researcher's decisions.
Frequently asked questions
- How big a gap between training and validation means overfitting?
- There is no universal threshold. What matters is whether reducing complexity improves the validation score. Some gap is normal — models almost always do better on data they have seen. The problem is a gap that widens as you add complexity while the validation score falls.
- Does 100% training accuracy always mean overfitting?
- Not always. Random forests and large networks often fit the training data perfectly and still generalize well. Overfitting is established only by the score on unseen data and by comparison with simpler variants.
- Does more data always cure overfitting?
- It usually helps, because noise in a larger sample cancels out more often and is harder to memorize. It will not help with underfitting — a model that is too simple stays too simple however many examples it gets. A learning curve shows which case you are dealing with.
Sources
- James G., Witten D., Hastie T., Tibshirani R. "An Introduction to Statistical Learning", 2nd ed., Springer 2021, ch. 2.2 (Assessing Model Accuracy).
- Hastie T., Tibshirani R., Friedman J. "The Elements of Statistical Learning", 2nd ed., Springer 2009, ch. 7 (Model Assessment and Selection).
- Goodfellow I., Bengio Y., Courville A. "Deep Learning", MIT Press 2016, ch. 5.2 (Capacity, Overfitting and Underfitting).
- Geman S., Bienenstock E., Doursat R. (1992). Neural Networks and the Bias/Variance Dilemma. Neural Computation, 4(1).
- scikit-learn documentation: Underfitting vs. Overfitting — https://scikit-learn.org/stable/auto_examples/model_selection/plot_underfitting_overfitting.html