ML Atlas

03 · Supervised · 5 min read · updated

What is machine learning and how is it different from ordinary programming?

In short

Machine learning means building programs that derive their rule from examples instead of being handed it by a programmer. What counts is accuracy on new data.

What it is

Machine learning (ML) is a field of computer science and statistics in which a program is not given a ready-made rule but derives one from examples. The algorithm picks, from some family of functions, the one that best maps inputs to the desired outputs, and success is judged by accuracy on data it has never seen.

Tom Mitchell's classic definition (1997) reads: a program learns from experience E with respect to task T and performance measure P if its performance at T, as measured by P, improves with E. The task might be spam detection, the experience thousands of labelled emails, and the measure the fraction of messages classified correctly.

The contrast is easiest to see on that same example. In traditional programming someone writes the rules: "if the subject contains 'free' and three exclamation marks, it is spam". In machine learning you give the algorithm labelled examples and it works out for itself which words point to spam, and how strongly. The rule comes from the data rather than from the programmer's head — which is why it can be built even where nobody is able to write it down, as in face recognition.

Mechanism — why it works this way

Almost every learning algorithm has three ingredients. The model is a family of functions with adjustable parameters: a straight line, a decision tree, a neural network. The loss function measures how far the predictions miss the truth. Optimization moves the parameters so that the loss on the training set goes down — for example by gradient descent or by greedy splitting of the data.

The catch is that we minimize the loss on training data while caring about future data. This works only under the assumption that future cases come from the same distribution as the training ones. When a hospital changes its diagnostic scanner or customers change their habits, a model trained on old data can quietly lose accuracy.

The second source of difficulty: examples alone never determine the rule uniquely. Infinitely many curves pass through any finite set of points. Every algorithm must therefore assume something about the world — that the relationship is linear, that similar cases have similar answers, that the rule is simple. This is called inductive bias. The "no free lunch" theorem says that, averaged over all possible problems, no algorithm beats any other; an advantage comes from matching the assumptions to the problem at hand.

This leads to the basic trade-off. A family of functions that is too rigid cannot capture the true relationship (underfitting). One that is too flexible memorizes the noise and the coincidences as well (overfitting). That is why data are split into a training part and a test part, and quality is judged only on the latter.

An important limitation: a model learns co-occurrences, not causes. If, in the training data, sick patients were more often imaged on one particular scanner, the model may learn to recognize the scanner instead of the disease.

By example

The Breast Cancer Wisconsin dataset contains 569 breast tumours described by 30 features of cell nuclei from biopsies (radius, texture, contour concavity and so on): 212 malignant and 357 benign. We split it at random (train_test_split, 25% for testing, stratified, random_state=0): 426 tumours for training, 143 for testing. A model that always says "benign" scores 62.9% accuracy on the test set — the baseline you must never fall below.

A single threshold rule chosen from the data ("worst nucleus perimeter above 106.1 → malignant") already reaches 88.8%. Logistic regression, which weighs all 30 features at once, achieves 95.8%: it is wrong in 6 cases out of 143 (3 missed malignant tumours, 3 false alarms). On the training data the same model scores 99.1% — the gap between those numbers is exactly the cost of generalizing to new cases. And one last lesson: the two kinds of error are not equally dangerous, so "accuracy" is only the beginning of an evaluation.

In practice

  • In scikit-learn every model shares the same interface: model.fit(X_train, y_train), then model.predict(X_test) or model.predict_proba(X_test).
  • Always start from a baseline: DummyClassifier (most frequent class) or DummyRegressor (mean).
  • Split the data before any preprocessing: train_test_split(..., stratify=y); wrap scaling and imputation in a Pipeline so no information leaks from the test set.
  • Use cross-validation (cross_val_score) to choose the model and hyperparameters, and open the test set once, at the very end.
  • Common mistake: reporting the score on training data, or tuning the model until it "works" on the test set.

Frequently asked questions

How is machine learning different from artificial intelligence?
Artificial intelligence is the broad goal: systems that perform tasks requiring intelligence. Machine learning is today's dominant way of reaching that goal — learning from data instead of hand-coding rules. Deep learning (neural networks with many layers) is in turn a subset of machine learning.
How is machine learning different from statistics?
The tools are largely shared (regression, probability, estimation). Statistics more often asks about inference — does an effect exist and how large is it — while machine learning asks about predictive accuracy on new data. The boundary is fluid, and many methods were born where the two fields meet.
How much data does a model need to learn?
There is no single number. It depends on the complexity of the relationship, the number of features and the noise level: linear regression with a few features is happy with hundreds of examples, while image recognition from scratch needs hundreds of thousands. In practice, plot a learning curve and check whether the validation score is still rising as you add examples.
Does a machine learning model "understand" the data?
Not in the human sense. A model finds statistical patterns that helped on the training data and does not know which of them are causal and which are accidental. That is why it is worth checking what a model bases its decisions on.

Sources

  • Mitchell T. "Machine Learning", McGraw-Hill, 1997, ch. 1.
  • James G., Witten D., Hastie T., Tibshirani R. "An Introduction to Statistical Learning", 2nd ed., 2021, ch. 2.
  • Hastie T., Tibshirani R., Friedman J. "The Elements of Statistical Learning", 2nd ed., 2009, ch. 2.
  • Géron A. "Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow", 3rd ed., 2022, ch. 1.
  • scikit-learn documentation, "Getting Started": https://scikit-learn.org/stable/getting_started.html

See also