ML Atlas

01 · Foundations · 5 min read · Interactive · updated

What is Bayes' theorem and how do you use it in practice?

In short

Bayes' theorem inverts a condition: from P(evidence|hypothesis) it gets P(hypothesis|evidence), combining the evidence with how common the hypothesis was.

What it is

Bayes' theorem lets you invert a conditional probability: P(H | D) = P(D | H) · P(H) / P(D). Here H is a hypothesis (e.g. "the patient is ill") and D is the data or evidence (e.g. "the test came back positive"). The formula comes from a paper by Thomas Bayes published posthumously in 1763; Pierre-Simon Laplace developed it in its general form.

The components have names. P(H) is the prior probability — how common the hypothesis was before we saw the evidence. P(D | H) is the likelihood — how well the hypothesis explains the evidence. P(H | D) is the posterior probability — our belief after taking the evidence into account. P(D) is the normalisation: the overall chance of seeing such evidence across all hypotheses.

The theorem is at once trivial and profound. Trivial, because it is two lines of algebra from the definition of conditional probability. Profound, because it describes how to learn rationally from data — and shows why human intuition so often goes wrong.

Mechanism — why it works this way

The derivation takes one line. The joint probability P(H and D) can be written in two ways: P(H) · P(D | H) or P(D) · P(H | D). Equating the two and dividing by P(D) gives Bayes' theorem. The denominator expands by the law of total probability: P(D) = P(D | H) · P(H) + P(D | not H) · P(not H).

The clearest version is the odds form: posterior odds = prior odds × likelihood ratio, where the likelihood ratio is P(D | H) / P(D | not H). Evidence multiplies the odds by a factor saying how many times more probable it is under one hypothesis than under the other. Strong evidence (a ratio of 100) for a rare hypothesis (odds of 1 to 1,000) still yields odds of only 1 to 10.

This is the source of the base rate fallacy. People judge P(H | D) by the strength of the evidence alone, ignoring P(H). A test that detects 90% of sick people seems "90% reliable", but if the disease is rare, most positive results come from healthy people, simply because there are far more of them. In the language of metrics: sensitivity is P(+ | ill), precision is P(ill | +), and Bayes' theorem links them through the frequency of the positive class.

The theorem also describes step-by-step learning: the posterior after the first piece of evidence becomes the prior for the next. Two independent positive tests mean multiplying the odds by the likelihood ratio twice. This idea underpins Bayesian inference, Kalman filters and belief updating in reinforcement learning.

In machine learning it is used directly by the naive Bayes classifier: P(class | features) ∝ P(class) · P(features | class), assuming the features are independent within each class. Caveat: the posterior is only as good as the prior and the likelihood. If the class frequency in the training data differs from that in production (e.g. after balancing the classes), the predicted probabilities will be systematically shifted.

By example

A simple numerical example: a disease affects 1% of people, the test detects 90% of those who are ill and gives a false alarm for 9% of those who are healthy. Out of 10,000 people there are 100 who are ill, of whom 90 test positive, and 9,900 who are healthy, of whom 891 test positive. P(ill | +) = 90 / (90 + 891) = 0.092. Despite a "good" test, only 9.2% of positive results are actually ill. A second, independent positive test raises this to about 50%.

On the Titanic dataset (891 passengers) you can check the theorem on real numbers. We know that P(survived | woman) = 0.742, P(woman) = 0.352 and P(survived) = 0.384. Bayes gives P(woman | survived) = 0.742 · 0.352 / 0.384 = 0.681, exactly the directly computed share of women among survivors (233 of 342). In odds form: the prior odds of survival are 342 : 549 = 0.62, and being a woman is 4.6 times more common among survivors than among victims. The posterior odds are 0.62 · 4.6 ≈ 2.88, i.e. a probability of 0.742.

In practice

  • sklearn.naive_bayes includes GaussianNB (continuous features), MultinomialNB (word counts) and BernoulliNB (binary features); set the prior with the priors or class_prior parameter.
  • When the class frequency in production differs from training, recompute the outputs with the odds form: multiply the odds by the ratio of the new to the old prior odds.
  • When multiplying many likelihoods, sum logarithms (predict_log_proba) — a product of hundreds of small numbers will round to zero.
  • You can estimate a model's precision at a different positive-class frequency from its sensitivity and specificity using Bayes' theorem.
  • Typical mistake: interpreting a p-value or a test's sensitivity as the probability that the hypothesis is true.

Frequently asked questions

Where does the prior probability come from?
From frequencies in the population (e.g. the prevalence of a disease), from earlier studies or from domain knowledge. In the absence of knowledge, weakly informative distributions are chosen; with plenty of data, the choice of prior has less and less influence on the result.
How does Bayesian statistics differ from frequentist statistics?
Both use Bayes' theorem when it comes to events. The difference concerns parameters: the Bayesian approach treats an unknown parameter as a random variable with a prior distribution, while the frequentist approach treats it as a constant, about which one reasons from how procedures behave over repeated samples.
Why does "naive" Bayes work when the features are not independent?
Because classification only requires the correct class to get the highest score, even if the probabilities themselves are exaggerated. The ranking of classes is often correct even though calibration is poor.

Sources

  • Bayes, 1763, "An Essay towards solving a Problem in the Doctrine of Chances", Philosophical Transactions of the Royal Society of London 53, 370–418.
  • Gigerenzer, Hoffrage, 1995, "How to improve Bayesian reasoning without instruction: Frequency formats", Psychological Review 102(4), 684–704.
  • Blitzstein, Hwang "Introduction to Probability", 2nd ed., CRC Press, 2019, ch. 2.
  • Bishop "Pattern Recognition and Machine Learning", Springer, 2006, ch. 1.2.
  • scikit-learn documentation: Naive Bayes, https://scikit-learn.org/stable/modules/naive_bayes.html

See also