ML Atlas

03 · Supervised · 5 min read · Interactive · updated

How does logistic regression work and how do you read its coefficients?

In short

Logistic regression is a classifier that turns a weighted sum of features into a probability with the sigmoid function. Its weights read as odds ratios.

What it is

Logistic regression is a binary classification model that predicts the probability of belonging to the positive class. It computes a weighted sum of features z = b₀ + b₁x₁ + … + bₚxₚ and passes it through the sigmoid function σ(z) = 1 / (1 + e^(−z)), which maps any number to a value between 0 and 1. Despite its name, it is a classifier, not a regression in the sense of predicting numbers.

The sigmoid has a gentle "S" shape: at z = 0 it gives 0.5, for large positive z it approaches 1, for large negative z it approaches 0. So the model says not only "yes/no" but also "how strongly yes". A decision is made only once you choose a threshold — 0.5 by default, but often lower in medicine or fraud detection.

The model was popularised in statistics by David Cox (1958). Today it is the standard baseline for classification, and a single neuron with a sigmoid is exactly logistic regression.

Mechanism — why it works this way

The key is odds: p / (1 − p). A probability of 0.75 corresponds to odds of 3 to 1. The logarithm of the odds, the logit, can take any value from minus to plus infinity, so it can be modelled with an ordinary weighted sum: log(p / (1 − p)) = b₀ + b₁x₁ + … The sigmoid is simply the inverse of the logit.

That gives the interpretation of the weights. Increasing feature xⱼ by one unit changes the logit by bⱼ, which multiplies the odds by e^(bⱼ). That number is the odds ratio: e^b = 2 means "twice the odds", e^b = 0.5 means "half the odds". The effect on the probability itself is not constant: the same jump in odds moves p a lot near 0.5 and barely at all when p is already close to 0 or 1.

The weights are fitted by maximum likelihood: we look for the parameters under which the observed labels are most probable. Equivalently, we minimise cross-entropy (log loss), which heavily penalises confident mistakes. This function is convex, so it has a single minimum, but there is no closed-form solution — it is solved iteratively (Newton's method, L-BFGS, gradient descent).

The decision boundary, the set of points where p = 0.5, is where z = 0 — a line, a plane or a hyperplane. The model is therefore linear in the features; curved boundaries require adding non-linear features.

Two caveats. When the classes are perfectly separable, the likelihood keeps growing as the weights grow, and without regularisation the coefficients run off to infinity. That is why scikit-learn applies an L2 penalty by default. Second, as in linear regression, collinearity makes individual weights unstable, and a weight is a conditional association, not a causal effect.

By example

Titanic: 891 passengers, 38.4% survived. Features: class (1–3), sex, age (177 missing values filled with the training-set median of 29 years), number of siblings/spouses and parents/children aboard, ticket fare. Training on 668 passengers, testing on 223 (stratified 75/25 split, random_state=0). The model is right on 77.1% of test cases, versus 61.4% for always guessing "did not survive".

The odds ratios tell the story of the disaster: being a woman multiplies the odds of survival by 14.4, each lower class multiplies them by 0.33, each year of age by 0.958 (a decade by 0.65), each sibling or spouse aboard by 0.69. In terms of probabilities: a 30-year-old woman travelling alone in 1st class on a £50 ticket — 94%; a 30-year-old man travelling alone in 3rd class on an £8 ticket — 9.5%. A curiosity: a model using sex alone predicts 18.5% for men and 74.9% for women — exactly the survival rates in the training set — and its odds ratio of 13.2 is the one computed directly from the 2×2 table. Here maximum likelihood simply reproduces the raw frequencies.

In practice

  • scikit-learn's LogisticRegression() uses an L2 penalty with C=1.0 by default (C is the inverse of the penalty strength); for unpenalised interpretation use penalty=None or statsmodels.Logit.
  • Scale features (StandardScaler) — it speeds up convergence and makes the penalty treat features equally.
  • Odds ratios: np.exp(model.coef_); probabilities: predict_proba, not predict.
  • Choose the decision threshold according to the costs of errors instead of leaving it at 0.5 by default; with imbalanced classes consider class_weight="balanced".
  • Multiple classes: scikit-learn uses multinomial logistic regression by default (softmax instead of the sigmoid).

Frequently asked questions

Why is it called "regression" if it is used for classification?
Because formally it models a number — the log odds — as a linear function of the features, just as linear regression models y. Classification only appears once a threshold is applied to the predicted probability. The name is historical and comes from statistics.
Are the probabilities from logistic regression reliable?
They are usually better calibrated than those of many other models, because the loss function directly rewards accurate probabilities. They are degraded by strong regularisation, a shift in class proportions between training and deployment, and missing non-linearities. It is worth checking a calibration curve.
How does logistic regression differ from a single neuron?
In no essential way: a neuron with a sigmoid and cross-entropy loss is logistic regression. A neural network stacks many such units into layers, which lets it form non-linear decision boundaries.

Sources

  • Cox D. R. "The Regression Analysis of Binary Sequences", Journal of the Royal Statistical Society B 20(2), 1958.
  • James G., Witten D., Hastie T., Tibshirani R. "An Introduction to Statistical Learning", 2nd ed., 2021, ch. 4.3.
  • Hastie T., Tibshirani R., Friedman J. "The Elements of Statistical Learning", 2nd ed., 2009, ch. 4.4.
  • Bishop C. "Pattern Recognition and Machine Learning", Springer, 2006, ch. 4.3.
  • scikit-learn documentation, "Logistic regression": https://scikit-learn.org/stable/modules/linear_model.html#logistic-regression

See also