English edition · 40 terms
Atlas of machine learning
How machine learning, neural networks and language models really work: every concept with the mechanism, an example on a well-known dataset and an illustration you can change yourself. The full atlas is in Polish; the English edition is growing.
English terms
Gradient and gradient descent InteractiveThe gradient is the vector of derivatives of the loss with respect to the weights; it points uphill. Gradient descent steps the weights the opposite way.Bayes' theorem InteractiveBayes' theorem inverts a condition: from P(evidence|hypothesis) it gets P(hypothesis|evidence), combining the evidence with how common the hypothesis was.What is machine learningMachine learning means building programs that derive their rule from examples instead of being handed it by a programmer. What counts is accuracy on new data.Linear regression InteractiveLinear regression predicts a number as a weighted sum of features plus a constant. Least squares picks the weights, and each one describes a feature's effect.Logistic regression InteractiveLogistic regression is a classifier that turns a weighted sum of features into a probability with the sigmoid function. Its weights read as odds ratios.k-nearest neighbors (kNN) InteractivekNN classifies a new point by a vote of its k nearest training examples. It is simple and flexible, but sensitive to the scale and number of features.Decision tree InteractiveA decision tree classifies an example with a series of yes/no questions about single features (age ≤ 6.5?) down to a leaf, cutting feature space into boxes.Random forest InteractiveA random forest averages hundreds of decision trees, each trained on a random data sample and random feature subsets. That makes it accurate and stable.Gradient boosting InteractiveGradient boosting builds a model step by step: each new shallow tree corrects the errors of the trees so far, moving in the direction that lowers the loss.Overfitting and underfitting InteractiveAn overfit model memorizes noise in the training set and does worse on new data; an underfit model is too simple to capture even the main pattern.Cross-validation InteractiveCross-validation splits the data into k parts and uses each in turn as the validation set. It gives a more stable estimate of performance than a single split.Confusion matrix InteractiveA confusion matrix counts how many examples of each true class the model assigned to each class. Accuracy, precision, recall and other metrics follow from it.Precision and recall InteractivePrecision is the share of a model's alarms that are real; recall is the share of real cases the model caught. Improving one usually makes the other worse.ROC AUC (area under the ROC curve) InteractiveROC AUC is the probability that a random positive example gets a higher score than a random negative one. It depends on the order of scores, not the threshold.K-means InteractiveK-means splits points into k groups: each point joins the nearest centre, and centres move to the mean of their group. It returns k groups even when none exist.Principal component analysis (PCA) InteractivePCA rotates the coordinate system so that the first axes capture as much of the data's variability as possible, replacing many correlated features with a few.Perceptron (artificial neuron)The simplest artificial neuron: it takes a weighted sum of inputs plus a bias and outputs 1 when it exceeds zero, splitting feature space with a hyperplane.Activation functions (ReLU, sigmoid, tanh) InteractiveAn activation function is the nonlinearity applied to a neuron's weighted sum. Without it a deep network is linear, and its shape decides how gradients flow.Multilayer perceptron (MLP) InteractiveAn MLP is a neural network with hidden layers that combine weighted sums and a nonlinearity, so it learns relationships a linear model cannot capture.Backpropagation InteractiveBackpropagation computes how the loss depends on every weight in a network by sending an error signal from the output back to the input via the chain rule.Learning rate InteractiveThe learning rate scales the step along the negative gradient when updating weights. Too small slows training; too large causes oscillation and divergence.Optimizers: SGD, momentum, Adam InteractiveAn optimizer turns a gradient into a step. SGD follows the gradient, momentum adds inertia, and Adam scales each weight's step by its typical gradient.Regularization (L2, dropout, early stopping) InteractiveRegularization stops a model from fitting noise: a penalty on large weights (L2), randomly switching off neurons (dropout) and ending training early.Dropout InteractiveDuring training, dropout randomly switches off some neurons, so the network cannot rely on individual connections and generalizes better to new data.Attention mechanism InteractiveAttention lets a model reach into any part of its input at every step and take as much as fits the current question — a weighted average by relevance.Self-attention InteractiveIn self-attention each token in a sequence forms a query, a key and a value, then gathers information from other tokens in proportion to how well they match.Transformer InteractiveThe transformer is an architecture of self-attention and MLP blocks with residual connections. It processes whole sequences at once and powers modern LLMs.Large language model (LLM) InteractiveAn LLM is a transformer neural network trained on a vast body of text to predict the next token. Its abilities grow out of that simple principle.Tokenization and BPE InteractiveTokenization splits text into tokens, chunks of words that a model turns into numbers. BPE builds the vocabulary by merging the most frequent symbol pairs.Word embeddings InteractiveA word embedding is a vector of numbers in which words used in similar contexts lie close together. It is the foundation on which an LLM builds meaning.Next-token prediction InteractiveFor every position in a text, an LLM computes a probability distribution over the next token. It learns via cross-entropy and writes by sampling token by token.Temperature and sampling in LLMs InteractiveTemperature divides the logits before the softmax: low values sharpen the distribution, high ones flatten it. Top-k and top-p cut off the unlikely tail.RAG — retrieval-augmented generation InteractiveRAG retrieves document passages that match a question and pastes them into the prompt. The model answers from sources rather than from training memory alone.Bias–variance tradeoff InteractiveA model's error on new data splits into bias, variance and noise. A simpler model has more bias, a richer one has more variance.Curse of dimensionality InteractiveAs dimensions grow, space empties out exponentially and distances between points become nearly equal. Neighbourhood-based methods stop making sense.Data leakage InteractiveData leakage is when a model, in training or validation, gets information that would not be available at real prediction time. It inflates the results.Goodhart's law InteractiveWhen a measure becomes a target, it ceases to be a good measure. In ML, optimising a proxy metric pulls it apart from what we actually care about.Central limit theorem InteractiveThe sum or mean of many independent values is approximately normally distributed, even when the individual values are extremely skewed. Here is why.Simpson's paradox InteractiveA relationship visible in the pooled data can vanish or reverse once the data is split into groups. The 1973 Berkeley admissions case, and why it happens.Anscombe's quartet InteractiveFour datasets with identical means, variances, correlation and regression line look completely different when plotted. Why summary statistics are not enough.