ML Atlas

04 · Evaluation · 4 min read · Interactive · updated

What is ROC AUC and why does the decision threshold not change it?

In short

ROC AUC is the probability that a random positive example gets a higher score than a random negative one. It depends on the order of scores, not the threshold.

What it is

ROC AUC (area under the ROC curve) is the area under the ROC curve, which for every possible decision threshold plots the share of positives caught (recall, TPR) against the share of false alarms (FPR). Numerically, AUC equals the probability that a randomly chosen positive example receives a higher score from the model than a randomly chosen negative one.

AUC = 1 means perfect ordering, 0.5 means random. It is the standard metric for binary classification, independent of the threshold and of class proportions, and widely used in competitions and in medicine.

Mechanism — why it works this way

The ROC curve is built like this: sort the examples by the model's score, sweep the threshold from the highest score to the lowest and at each point record the pair (FPR, TPR). Each positive example moves the curve up, each negative one moves it right. The area beneath it is the share of (positive, negative) pairs in which the positive ranks higher; ties count as half. Hence the equivalence with the Mann–Whitney U statistic: AUC = U / (n₊ · n₋) (Hanley and McNeil 1982).

A consequence that surprises people: there is no threshold in the formula. AUC evaluates the ordering (the ranking), not the decisions. Moving the threshold changes recall, precision, accuracy and balanced accuracy, because those evaluate decisions at a single cut, but it does not change AUC one bit. Likewise, any strictly increasing transformation of the scores — temperature scaling, a logarithm, Platt calibration — preserves the order and therefore preserves AUC. A constant score for every example produces nothing but ties, i.e. AUC = 0.5.

AUC also does not depend on class proportions, because it counts pairs, not rows. This is an advantage when comparing models, but also a trap: with a very rare positive class AUC can be high while precision is dismal, because FPR, computed relative to the huge negative class, stays small even with many false alarms. In that case the precision–recall curve and the area under it (average precision) are a better choice.

Caveat: AUC measures the quality of the ranking, not the quality of the probabilities — a model with an AUC of 0.95 can be terribly calibrated. If you ultimately make decisions at a single threshold, AUC is only an indirect indicator.

By example

Breast Cancer Wisconsin, tested on 171 tumours (64 malignant, random_state=0). A deliberately weak model — logistic regression on just the mean texture and mean smoothness of the tumour — has an AUC of 0.77; the formula U / (n₊ · n₋) computed with scipy.stats.mannwhitneyu gives the same result. At a threshold of 0.2 the model catches 83% of malignant tumours with 49% false alarms (BA 0.67), at 0.3 — 78% with 38% (BA 0.70), at 0.5 — 53% with 19% (BA 0.67), and at 0.7 — 30% with 5% (BA 0.63). Balanced accuracy, recall and FPR change with the threshold; AUC stays at 0.77 throughout.

Dividing the logits by a temperature of 0.5, 2 or 5 changed the mean predicted probability from 0.32 to 0.46, while AUC remained exactly 0.772. A constant score for all tumours gives 0.5, and the full model on 30 features gives 0.99.

In practice

  • scikit-learn: roc_auc_score(y_true, scores) — pass probabilities or continuous scores, not hard labels from predict(), which leave just one interior point on the curve; roc_curve, RocCurveDisplay.
  • Multiclass: roc_auc_score(..., multi_class="ovr") or "ovo" with averaging.
  • XGBoost/LightGBM: eval_metric="auc" for early stopping.
  • An AUC below 0.5 means the model orders things backwards — most often a sign error or swapped labels.
  • Typical mistake: tuning the threshold "for AUC" — the threshold has no effect on AUC; it is tuned for decision metrics.

Frequently asked questions

How do you interpret an AUC of 0.8?
If you draw one positive and one negative example at random, in 80% of cases the model will give the positive one a higher score. 0.5 is guessing, 1.0 is perfect ordering. The value describes the ranking, not how many decisions at a chosen threshold will be correct.
Why doesn't AUC change when the threshold changes?
Because AUC is computed from all thresholds at once — it is the area under the curve that the threshold traces out. Choosing a threshold means choosing a point on that curve; the curve and the area beneath it stay put. Changing the threshold changes recall, precision and accuracy, not AUC.
ROC AUC or average precision?
With reasonably balanced classes, ROC AUC is the standard and easy to interpret. With a very rare positive class, ROC AUC can be optimistic; the precision–recall curve then shows the true cost of false alarms.

Sources

  • Fawcett, T. (2006). "An introduction to ROC analysis". Pattern Recognition Letters 27(8), 861–874. doi:10.1016/j.patrec.2005.10.010
  • Hanley, J. A., McNeil, B. J. (1982). "The meaning and use of the area under a receiver operating characteristic (ROC) curve". Radiology 143(1), 29–36.
  • Davis, J., Goadrich, M. (2006). "The relationship between precision-recall and ROC curves". ICML.
  • Géron, A. (2022). Hands-On Machine Learning, 3rd ed., O'Reilly, ch. 3 "Classification" (the ROC curve).
  • scikit-learn: "Receiver operating characteristic (ROC)". https://scikit-learn.org/stable/modules/model_evaluation.html#roc-metrics

See also