011 min
The problem large language models solve
The approach that came before was pre-train then fine-tune, and it worked. A model was pre-trained on general text, then fine-tuned on a labelled dataset for one specific task: sentiment classification, named entity recognition, question answering.
The cost was in the second step. Brown and colleagues state it precisely in the 2020 GPT-3 paper: while the architecture was task-agnostic, the method still required task-specific fine-tuning datasets of thousands or tens of thousands of examples. Every new task meant collecting and labelling a new dataset.
That is a bottleneck with a particular shape. It is not that the models were bad. It is that the unit of work was the task. A company with forty text-processing problems needed forty labelled datasets, forty training runs and forty deployed models, and any problem too small to justify a labelling budget simply did not get solved.
The paper frames the target by contrast with people: humans can generally perform a new language task from a few examples or from simple instructions, which NLP systems at the time largely could not do.
So the problem is not language understanding in the abstract. It is the cost of the second step, and whether a single model can be made general enough that the second step is a few sentences of instruction instead of ten thousand labelled rows.
022 min
How a large language model is built
Follow one thing the whole way through: a model that can answer what is the capital of Australia?
Pretraining: predict the next token
The model is shown enormous quantities of text and given exactly one job. Given everything so far, predict what comes next. Text is split into tokens, which are common chunks of characters rather than whole words. The model outputs a probability for every token in its vocabulary, is scored on how much probability it gave the one that actually followed, and its weights are nudged to do better.
That is the entire training objective. There is no step where anyone tells it facts about Australia.
Why does geography come out of it anyway? Because the text contains sentences like Canberra, the capital of Australia, has a population of and thousands of variants. To predict the token after the capital of Australia is, a network has to encode something about the relationship between those words. Prediction accuracy on a large enough corpus cannot be achieved by surface statistics alone, so the pressure to predict well becomes pressure to represent grammar, facts, and patterns of argument.
This is the claim worth being careful about. The model was optimised for prediction. Everything else is a by-product that turned out to be load-bearing.
What makes it large
Two numbers grow together. GPT-3 had 175 billion parameters, which the paper notes was ten times more than any previous non-sparse language model, and it was trained on a correspondingly large corpus.
Post-training: make it useful
A pretrained model completes text. Asked what is the capital of Australia?, it may well produce three more quiz questions, because that is what follows a question in a lot of documents. It is doing its job correctly and is not doing yours.
The fix has two stages, set out by Ouyang and colleagues in 2022. First, people write demonstrations of the desired behaviour and the model is fine-tuned on them by supervised learning. Second, people rank model outputs against each other, those rankings train a reward model, and the language model is further tuned by reinforcement learning against it. That second stage is reinforcement learning from human feedback.
The result is a model that answers Canberra β not because it learned a new fact, but because it learned which of its many plausible continuations is the one a person asking that question wants.
031 min
A concrete example: when small beat large
The clearest evidence that post-training is not a cosmetic layer is a single comparison from the InstructGPT paper.
Ouyang and colleagues took GPT-3 and applied the two-stage process above. They then ran human evaluations comparing the outputs of the tuned model against the original.
Outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters.
A model with one hundred times fewer parameters produced answers people preferred. The paper opens by stating the reason directly: making language models bigger does not inherently make them better at following a user's intent, and large models can produce outputs that are untruthful, toxic, or simply unhelpful.
Two things follow that are easy to get wrong in the other direction.
The smaller model did not become more knowledgeable. Capability from pretraining is roughly what it was; what changed is how reliably that capability is aimed at what was asked. Preference ratings measure helpfulness, not accuracy.
And this result does not say scale is unimportant. It says scale and instruction-following are different axes, and that a product experience most users would call quality lives substantially on the second one.
042 min
What a large language model costs to build
For years the assumption was that a fixed compute budget is best spent on more parameters. Hoffmann and colleagues tested it in 2022 and found the field had been doing it wrong.
They trained over 400 models, from 70 million to more than 16 billion parameters, on 5 to 500 billion tokens. Their conclusion is stated in one line in the abstract: for compute-optimal training, model size and the number of training tokens should be scaled equally, so for every doubling of model size the training data should double too.
The headline consequence is that current large models were significantly undertrained. Their demonstration model, Chinchilla, used 70 billion parameters against the 280 billion of Gopher, trained on four times the data for the same compute budget.
Three practical readings follow.
Parameter count is a weak proxy for capability. Comparing two models by size alone ignores the axis that turned out to matter at least as much.
Data is a binding constraint, not an abundant input. If tokens must scale with parameters, high-quality text becomes the scarce resource, which is why data curation and licensing became strategic rather than operational concerns.
Inference cost is separate from training cost, and it is the one you pay. A smaller model trained on more data is cheaper to serve for its capability, which matters far more to anyone deploying than the training bill they never see.
052 min
What this means for your work
For a product manager: the model is not the product surface. Two teams using the same model ship very different quality, and the difference is in instructions, retrieval, tool access and evaluation. Budget for those as the work rather than as configuration.
For an engineer: the training objective explains the failure modes. A model optimised to produce likely continuations will produce a fluent, likely-looking answer when it has no grounding for one. That is the mechanism behind fabricated citations and invented function names, and it is why the fix is architectural β retrieval, tool use, verification β rather than a better-worded instruction not to make things up.
For an engineer, second: nothing persists between calls. A model has no memory of your last request. Anything it appears to remember was re-sent by your application. This is covered in the context window entry and it is the most common source of surprise in a first production deployment.
For a designer: outputs vary between identical requests. Sampling makes the same prompt produce different text. Interfaces that assume one canonical answer β a cached result, a shareable permalink, a diff against last time β need an explicit decision about what gets stored, because the model will not reproduce it on request.
For a founder: capability claims decay fast, so date them. A published benchmark number describes one model version on one dataset on one date. As of September 2026 the field has changed materially several times since GPT-3, and any plan resting on what models cannot currently do needs an explicit review date attached.
062 min
What large language models do not solve
Anything requiring a guarantee. The output is sampled from a distribution. If a wrong answer is unacceptable rather than merely unwelcome, the model belongs behind something that checks: a validator, a database lookup, a human, a type system. The right tool for a task with a correct answer that can be computed is the computation.
Knowing what it does not know. The same objective that produces fluent correct answers produces fluent wrong ones, and the model's confidence is not a reliable signal of which you have. Brown and colleagues flagged limitations of this kind in the GPT-3 paper itself, naming datasets where few-shot learning struggled and methodological issues arising from training on large web corpora.
Facts after its training cutoff, or facts that were never public. This is not a defect to be prompted around. It is the reason retrieval exists, and the retrieval-augmented generation entry covers the pattern.
The well-known overclaim, named and corrected. "It understands" is doing unmarked work in most sentences that contain it. What is established is that these systems produce text that is frequently correct and useful across a very wide range of tasks. Whether that constitutes understanding is a live argument, not a settled premise, and a product decision that rests on the strong reading is resting on the contested part.
And the cheerful version of the same error: a demo that works is evidence about the demo. Without an evaluation set drawn from your own data, a convincing first session tells you almost nothing about the hundredth.
071 min
Large language model vs. nearby concepts
| Compared with | The one fact that decides |
|---|---|
| AITransformer | |
| AIChatbot | |
| Not in the library yetArtificial general intelligence | |
| AIRetrieval-augmented generation |
When the question is "should we use an LLM or fine-tune one", it is malformed, because a fine-tuned model is still one. The real question is usually whether to change the weights or change the input, and the second is cheaper, faster to revise and easier to evaluate.
082 min
Where the evidence is contested
The sharpest live disagreement is about whether scaling produces genuinely new capabilities, and it is worth following because it determines how much weight to put on forecasts.
Jason Wei and sixteen co-authors made the case for emergence in 2022. Their definition is precise: an ability is emergent if it is not present in smaller models but is present in larger ones. The consequence they draw is the interesting part β such abilities cannot be predicted by extrapolating from smaller models, so further scaling could keep expanding what these systems can do in ways nobody can forecast.
Rylan Schaeffer, Brando Miranda and Sanmi Koyejo published the rebuttal in 2023 under the title Are Emergent Abilities of Large Language Models a Mirage? Their claim is that emergent abilities appear because of the researcher's choice of metric rather than a fundamental change in model behaviour with scale. A metric that awards credit only for an exactly correct answer turns smooth underlying improvement into an apparent jump. Under a linear metric, the same models improve gradually and predictably.
They support this three ways: confirming predictions about metric effects on tasks with claimed emergent abilities, a meta-analysis across BIG-Bench, and a demonstration in which they deliberately produced never-before-seen apparently emergent abilities in vision models purely by choosing metrics. Their conclusion is that alleged emergent abilities evaporate under different metrics or better statistics.
Stated fairly, this is not a claim that large models cannot do things small ones cannot. It is a claim about whether the transition is a discontinuity in the system or an artefact of the ruler.
Where this leaves a practitioner is concrete. Be sceptical of roadmaps premised on unpredictable capability jumps, since the strongest evidence for unpredictability is contested. And choose evaluation metrics deliberately: if an all-or-nothing metric can manufacture the appearance of a sudden jump, it can equally hide steady progress on your own task. Graded scoring will tell you more about whether last month's work helped than a pass-rate will.
091 min
How the large language model changed since
The term has stayed put while almost everything it names has moved.
2020 β scale as the thesis. GPT-3 established that a single model at 175 billion parameters could handle many tasks from instructions alone, without gradient updates. The argument was that generality comes from size.
2020 to 2022 β the recipe corrected. Kaplan and colleagues published scaling laws relating performance to model size, data and compute. Hoffmann and colleagues then showed in 2022 that the prevailing reading had over-weighted parameters and under-weighted data, and that model size and training tokens should scale together.
2022 β the second half of training arrives. InstructGPT showed that a 1.3-billion-parameter model tuned on human feedback was preferred to the 175-billion-parameter original. Post-training stopped being an afterthought and became a stage with its own research programme.
2022 to 2023 β the emergence argument. Claims of unpredictable capability jumps were made and substantially contested, as above.
What this means for reading anything written about these systems: the phrase means something different depending on when it was written. A 2020 description is about a pretrained text completer. A 2026 description is about a pretrained model plus a post-training pipeline, usually with tool access and retrieval attached. The label did not change and the object did.
?8 questions
Questions people ask
What does the large in large language model refer to?
Is a large language model the same as a Transformer?
How does an LLM learn facts if it is only predicting text?
What is RLHF and why is it needed?
Does a bigger model always perform better?
Why do large language models make things up?
Do large language models have emergent abilities?
Should I fine-tune a model or change the prompt?
Β§7 sources
Sources
Brown, T. B. et al. (2020). Language Models are Few-Shot Learners (GPT-3). arXiv:2005.14165
Ouyang, L. et al. (2022). Training language models to follow instructions with human feedback (InstructGPT). arXiv:2203.02155
Hoffmann, J. et al. (2022). Training Compute-Optimal Large Language Models (Chinchilla). arXiv:2203.15556
Kaplan, J. et al. (2020). Scaling Laws for Neural Language Models. arXiv:2001.08361
Show all 7 sourcesShow fewer sources
Wei, J. et al. (2022). Emergent Abilities of Large Language Models. arXiv:2206.07682
Schaeffer, R., Miranda, B. and Koyejo, S. (2023). Are Emergent Abilities of Large Language Models a Mirage? arXiv:2304.15004
Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762





