08 · LLMs · 5 min read · Interactive · updated
What is a large language model (LLM) and how does it work?
In short
An LLM is a transformer neural network trained on a vast body of text to predict the next token. Its abilities grow out of that simple principle.
What it is
A large language model (LLM) is a neural network, usually with a transformer architecture, that computes, for a given piece of text, a probability distribution over the next token (a chunk of a word). It has anywhere from a few billion to hundreds of billions of parameters and is trained on trillions of tokens of text. It generates a response by repeatedly sampling the next token and appending it to the input.
Intuition: imagine your phone's autocomplete having read a large part of the public internet, books and code. To guess well how sentences like "The capital of France is…", "def fibonacci(n):" or "The conclusion that follows from these premises is…" continue, the model has to learn grammar, facts, programming conventions and common patterns of reasoning. Nobody teaches it these things separately — they emerge because they help predict text.
The word "large" is a matter of convention. In practice it means models big enough to do things they were not explicitly trained for: translate, summarize, write code, solve tasks from a few examples given in the prompt.
Mechanism — why it works this way
From text to numbers. Text is first split into tokens (see BPE tokenization), and each token is turned into a vector — an embedding. It then passes through a stack of identical transformer blocks. In each block, the self-attention mechanism lets every position "look at" earlier positions and pull in the information it needs, while a fully connected layer (MLP) processes each position separately. At the end, the vector at the last position is projected onto a vector of scores (logits), one for every token in the vocabulary, which the softmax function turns into probabilities.
There is a single training objective. In pretraining the model sees a piece of text and has to predict every next token. The loss function is cross-entropy: a penalty equal to −log of the probability the model assigned to the true next token. Gradient descent moves billions of weights so that this penalty decreases on average. Because the label is simply the continuation of the text, no manual labelling is needed — this is self-supervised learning.
Why "understanding" grows out of prediction. Good prediction requires compression: a model that knows a rule needs less "memory" than one that memorizes cases. The cheapest way to reach a low loss on diverse text turns out to be an internal representation of syntax, meaning and the causal relationships described in the text. Kaplan et al. (2020) showed that the loss falls predictably, as a power law, with the number of parameters, the amount of data and compute — hence the race for scale.
Base model versus assistant. A model after pretraining alone continues text but does not follow instructions: asked a question, it may reply with another question. A chat assistant is obtained by fine-tuning on examples of instructions and responses and by learning from human preferences (RLHF, DPO).
Limitations built into the design. An LLM models the statistics of text, not the world directly. It has no built-in truth checking, so it fluently generates false statements too (hallucinations). Its knowledge is frozen at the moment training ended. The model has no memory between conversations — it sees only what fits in its context window. It generates token by token, with no way to take back a word it has already written.
By example
Where does the figure of "175 billion parameters" in GPT-3 come from? The model has 96 blocks of width d = 12,288 (Brown et al., 2020). One transformer block has about 12·d² weights: 4·d² in self-attention (the query, key, value and output matrices, each d×d) and 8·d² in the MLP layer (two d×4d matrices). That gives 12 · 96 · 12,288² ≈ 173.9 billion. The embedding table for a vocabulary of 50,257 tokens adds 50,257 · 12,288 ≈ 0.62 billion. Together that is about 174.6 billion, i.e. the stated 175 billion (the rest consists of small bias and normalization vectors).
The weights alone, stored in 16-bit precision (2 bytes per parameter), take up 175 billion · 2 B = 350 GB. That is why the largest models cannot run on a single graphics card, and why quantization and distillation matter so much in practice.
In practice
- Open models are usually run through the
transformerslibrary (AutoModelForCausalLM,AutoTokenizer,model.generate) or lighter inference engines with quantization. - Match the model size to the task: models with 1–8 billion parameters are enough for classification, extraction and simple conversation; hard reasoning and code benefit from larger ones.
- Before fine-tuning, try a good prompt and RAG — they are cheaper and easier to change.
- Cost and latency are counted in tokens (input and output), not in characters or words.
- Common mistake: treating an answer like a database lookup. Facts, numbers and quotations must be verified.
Frequently asked questions
- Does an LLM understand what it writes?
- The model has internal representations that let it use concepts correctly in new contexts, so in a functional sense it "understands" something. It has no access to the world beyond text, however, and no mechanism for telling truth from plausible-sounding falsehood. The philosophical debate continues; what matters in practice is that it can be confident even when it is wrong.
- How is an LLM different from an ordinary neural network?
- The principle is the same: layers, weights, gradient descent. What differs is the scale, the architecture (a transformer with self-attention) and the training objective (next-token prediction on unlabelled text). That combination produces a general-purpose model instead of a single-task one.
- How does an LLM know things that were not in its data?
- Usually it doesn't — it combines patterns that were in the data. It can apply a rule correctly to a new case, but new facts from after its training cutoff have to be supplied in the context, for example through RAG or tools.
Sources
- Vaswani A. et al., 2017, "Attention Is All You Need", NeurIPS 2017.
- Radford A. et al., 2019, "Language Models are Unsupervised Multitask Learners", OpenAI technical report.
- Brown T. et al., 2020, "Language Models are Few-Shot Learners", NeurIPS 2020.
- Kaplan J. et al., 2020, "Scaling Laws for Neural Language Models", arXiv:2001.08361.
- Jurafsky D., Martin J. H., "Speech and Language Processing", 3rd ed. (online draft), chapters on large language models and transformers.