ML Atlas

07 · Architectures · 5 min read · Interactive · updated

What is the attention mechanism in neural networks and how does it work?

In short

Attention lets a model reach into any part of its input at every step and take as much as fits the current question — a weighted average by relevance.

What it is

Attention is an operation in which a model, for the current "question" (the query), scores how well each available element (the keys) matches it, turns those scores into weights that sum to 1, and returns a weighted average of the content associated with them (the values). The model therefore does not have to squeeze the whole input into a single vector — at every step it can "look" wherever the information it needs happens to be.

Intuition: it is a soft dictionary lookup. An ordinary dictionary returns the value for the key that matches exactly. Attention returns a mixture of all the values, with more of those whose keys resemble the query. When translating the Polish sentence "Kot pije mleko" ("The cat drinks milk") into English, the model puts the greatest weight on "mleko" while generating the word "milk", even though technically it sees all the words.

The mechanism was introduced by Bahdanau, Cho and Bengio (2015) for RNN-based machine translation. In 2017 Vaswani and colleagues showed that an entire model can be built from attention alone — this is how the transformer, the foundation of modern language models, came about.

Mechanism — why it works this way

The most popular version is scaled dot-product attention: Attention(Q, K, V) = softmax(Q·Kᵀ / √d_k) · V. Step by step: (1) the dot product of the query with each key measures how well they match — it is large when the vectors point in a similar direction; (2) dividing by √d_k keeps the scores on a reasonable scale; (3) softmax turns the scores into positive weights that sum to 1; (4) the result is the sum of the values weighted by these weights.

Why does this solve the RNN problem? The classic encoder–decoder squeezed the whole source sentence into a single state vector, which became a bottleneck for long sentences — translation quality fell with length. Attention gives the decoder direct access to the state of every input word. The path between any two positions has length 1, so the gradient does not have to pass through dozens of steps.

Why divide by √d_k? If the components of the vectors q and k are random with mean 0 and variance 1, their dot product has variance d_k, i.e. a standard deviation of √d_k. With d_k = 64, the scores therefore spread out with a standard deviation of about 8. A softmax of such large numbers produces almost all-or-nothing weights, and then its gradients are close to zero and training stalls. Dividing by √d_k brings the standard deviation back to about 1.

Attention weights are data-dependent — computed afresh for every input. This is what distinguishes attention from an ordinary dense layer, whose connection weights are fixed after training. Attention is also differentiable, so the model learns, by backpropagation, how to formulate its queries and keys.

Caveats: the cost of computing all query–key pairs grows quadratically with sequence length. Attention weights are tempting as an "explanation" of a model's decision, but research shows they do not always indicate what actually influenced the output — they should be treated with caution.

By example

Three tokens, two-dimensional vectors (d_k = 2). Query q = [2, 1]. Keys: k_1 = [1, 0], k_2 = [1, 1], k_3 = [0, −1]. Values: v_1 = [10, 0], v_2 = [0, 10], v_3 = [5, 5]. The dot products q·k are 2, 3 and −1. Divided by √2 ≈ 1.414: 1.41, 2.12 and −0.71. The softmax gives weights of 0.32, 0.64 and 0.04 (summing to 1). Result: 0.32·[10, 0] + 0.64·[0, 10] + 0.04·[5, 5] ≈ [3.37, 6.63]. The query "looks" mainly at the second token, a little at the first, and almost ignores the third, whose key points in the opposite direction.

Scaling matters. If the same scores were four times larger (8, 12, −4), as happens in higher dimensions without dividing by √d_k, the softmax would give weights of 0.02, 0.98 and 0.00 — an almost hard choice of one token and a near-zero gradient for the others. We also checked the scaling rule: for 100,000 pairs of random 64-dimensional vectors with components drawn from N(0, 1), the standard deviation of the dot product was 8.0, and after dividing by √64 = 8 it was 1.0.

In practice

  • PyTorch: torch.nn.functional.scaled_dot_product_attention(q, k, v, is_causal=True) — an efficient implementation (including FlashAttention) with a causal mask for generative models.
  • A ready-made layer with projections and multiple heads: nn.MultiheadAttention(embed_dim, num_heads, batch_first=True).
  • A causal mask forbids looking at future tokens — without it a language model "cheats" by seeing the answer during training.
  • A padding mask (key_padding_mask) excludes artificial padding tokens from the attention weights.
  • Visualizing attention weights (heat maps) helps with debugging, but it is not a full explanation of the model's decision.

Frequently asked questions

What are queries, keys and values?
They are three roles played by the same data. The query describes what we are looking for; the key, how an element "advertises" itself; the value, what content it hands over if it is selected. In models these are usually three different linear projections of the same vectors.
How is attention different from self-attention?
In cross-attention the queries come from one sequence (e.g. the translation being produced), and the keys and values from another (the source sentence). In self-attention all three come from the same sequence, so each token looks at the other tokens of the same text.
Do attention weights explain why a model made a decision?
Only partly. They show where the model took information from in a given layer, but the model has many layers and heads, and other weight configurations can lead to the same answer. Attribution methods are better suited to explanation.

Sources

  • Bahdanau, Cho, Bengio "Neural Machine Translation by Jointly Learning to Align and Translate", ICLR 2015, arXiv:1409.0473.
  • Vaswani et al. "Attention Is All You Need", NeurIPS 2017, arXiv:1706.03762.
  • Jain, Wallace "Attention is not Explanation", NAACL 2019.
  • Zhang et al. "Dive into Deep Learning", d2l.ai, ch. 11 ("Attention Mechanisms and Transformers").
  • PyTorch documentation: torch.nn.functional.scaled_dot_product_attention, https://pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention.html

See also