ML Atlas

05 · Unsupervised · 4 min read · Interactive · updated

What is principal component analysis (PCA) and what is it used for?

In short

PCA rotates the coordinate system so that the first axes capture as much of the data's variability as possible, replacing many correlated features with a few.

What it is

Principal component analysis (PCA) is a dimensionality reduction method that replaces the original, often correlated features with new axes — the principal components. The first component is the direction along which the data is most spread out; each subsequent one is the direction of greatest spread among those perpendicular to the previous ones. By keeping the first few components, you lose as little information about the variability of the data as possible.

Intuition: a point cloud shaped like a flattened cigar in three dimensions. The axis along the cigar tells you the most, the axis across its width a little, and its thickness almost nothing. PCA finds these axes automatically and then lets you "forget" about the thickness. Each component is a weighted sum of the original features, and the weights (loadings) tell you what the axis measures.

Mechanism — why it works this way

Projecting the data onto a direction given by a unit vector w yields a variance of wᵀΣw, where Σ is the covariance matrix (after centring the data). We look for the w that maximises this variance. The solution is the eigenvector of Σ with the largest eigenvalue, and the eigenvalue itself is the variance along that axis. Subsequent components are the subsequent eigenvectors. Because Σ is symmetric, its eigenvectors are perpendicular, and the new coordinates are uncorrelated.

There is a second, equivalent view: PCA with k components gives the best approximation of the data by a k-dimensional subspace in the least-squares sense. Maximising the variance of the projection and minimising the reconstruction error are the same thing, because total variance = retained variance + error (Pythagoras' theorem for each point). In practice PCA is computed via the SVD of the data matrix, without forming Σ explicitly.

The proportion of variance explained by a component is its eigenvalue divided by the sum of all eigenvalues. A plot of these proportions (a scree plot) helps choose the number of components: usually enough to explain 80–95% of the variance in total, or up to the point where the curve flattens out.

Two caveats are fundamental. First, PCA is sensitive to scale — a feature measured in large units has a large variance and will dominate the first component, which is why features are usually standardised (equivalent to PCA on the correlation matrix). Second, PCA looks only for linear directions and maximises variance, not usefulness for any particular task. High variance does not necessarily mean "important information".

By example

The Wine dataset: 178 wines from three grape cultivars, 13 chemical features. After standardisation, the first component explains 36.2% of the variance, the second 19.2%, together 55.4%; three components — 66.5%. Reaching 90% takes 8 of the 13 components, reaching 95% takes 10. Only the first three eigenvalues exceed 1 (4.73, 2.51, 1.45), meaning only they carry more variability than a single standardised feature.

The loadings let you name the axes. The first component has large positive weights on flavanoids (0.42), total phenols (0.39), the OD280/OD315 ratio (0.38) and proanthocyanins (0.31), and negative ones on nonflavanoid phenols (−0.30) and malic acid (−0.25) — a "polyphenol axis". The second is mainly colour intensity (0.53), alcohol (0.48) and proline (0.36). Logistic regression (5-fold cross-validation) identifies the cultivar correctly in 98.3% of cases using all 13 features, 95.5% using two components and 84.9% using one. Two numbers instead of thirteen cost about 3 percentage points of accuracy.

In practice

  • In scikit-learn: make_pipeline(StandardScaler(), PCA(n_components=2)); variance proportions in explained_variance_ratio_, loadings in components_.
  • PCA(n_components=0.95) picks the number of components that explain 95% of the variance by itself.
  • For very large data: IncrementalPCA or PCA(svd_solver="randomized"); for sparse matrices (e.g. TF-IDF) TruncatedSVD.
  • Fit PCA on the training set only — inside a pipeline, together with the model — to avoid data leakage during validation.
  • inverse_transform reconstructs the data from a few components; the reconstruction error is sometimes used for anomaly detection.
  • Typical uses: 2D visualisation, compression, removing collinearity before regression, denoising.

Frequently asked questions

How many principal components should you keep?
Popular rules: enough to explain 80–95% of the variance; the "elbow" in the variance plot; components with an eigenvalue above 1 for standardised data. If PCA is a step before a model, it is best to treat the number of components as a hyperparameter and choose it by cross-validation.
Do you need to standardise data before PCA?
Almost always when features are in different units — otherwise the first component will simply reflect the feature with the largest numbers. You can skip standardisation only when all features share the same unit and their variances are naturally comparable, e.g. image pixels.
Does PCA select the most important features?
No. PCA creates new features as combinations of all the old ones; it does not pick a subset of the original columns. Feature selection methods do that; PCA loadings only tell you which features contribute to the directions of high variability.

Sources

  • Pearson K., "On Lines and Planes of Closest Fit to Systems of Points in Space", Philosophical Magazine 2(11), 1901.
  • Hotelling H., "Analysis of a Complex of Statistical Variables into Principal Components", Journal of Educational Psychology 24, 1933.
  • Jolliffe I. T., Cadima J., "Principal component analysis: a review and recent developments", Philosophical Transactions of the Royal Society A 374, 2016.
  • James G., Witten D., Hastie T., Tibshirani R., "An Introduction to Statistical Learning", 2nd ed., Springer 2021, ch. 12.2 (Principal Components Analysis).
  • scikit-learn documentation, "PCA": https://scikit-learn.org/stable/modules/decomposition.html#pca

See also