011 min
Key takeaways
- What it is: prompting a model to show its reasoning steps before its final answer
- Why it exists: a direct answer skips steps a model needs to get multi-step problems right
- What it costs: more output tokens and higher latency for every answer
- When it breaks: the gains largely appear only in sufficiently large models
- As of: the original 2022 papers reported the largest gains on math and logic benchmarks
021 min
The problem Chain-of-Thought Prompting solves
Before chain-of-thought prompting, the standard way to get a large language model to answer a question was to show it a few examples of a question paired directly with its final answer, with nothing in between, and then ask a new question the same way. That worked well for a task a model could answer in a single step, like classifying a sentence's sentiment, but it failed on a task that needs several steps of reasoning chained together, like a multi-step arithmetic word problem. Asked directly for the final number, a model would frequently skip a step or combine two numbers incorrectly, the same way a person asked to do multi-digit arithmetic entirely in their head, with no scratch paper, makes more errors than the same person working it out on paper. The model had no equivalent of scratch paper: standard prompting gave it no space to work anything out before it had to commit to a final answer. Simply making the model bigger did not fix this on its own; a larger model prompted the same direct way still skipped the same intermediate steps, which suggested the problem was in how the question was asked, not only in how much the model knew.
032 min
How Chain-of-Thought Prompting works
Chain-of-thought prompting changes what the examples in a prompt look like, not the model itself. Instead of pairing each example question directly with its final answer, the prompt pairs each example question with a short written explanation of the reasoning that leads to the answer, and only then the answer itself. When the model is later given a new question in the same format, it continues the pattern: writing out its own reasoning steps before producing a final answer.
The mechanism works because a large language model, almost always built on the Transformer architecture, generates its response one token at a time, each new token conditioned on every token that came before it, including the tokens it has just generated itself. Writing an intermediate reasoning step first means every later token, including the final answer, is generated while conditioned on that reasoning text sitting in front of it, functioning like a working memory the model can refer back to. Skipping straight to the answer removes that scratch space entirely: the model has to arrive at the right answer using only its internal computation for a single answer token, with no intermediate text to condition on. A version of the same idea needs no worked examples at all. Kojima and colleagues showed that simply adding the words "let's think step by step" before a model's answer, with no example reasoning shown at all, triggers the same behaviour, because instructive language alone is often enough to make the model produce reasoning text instead of jumping to an answer.
041 min
A concrete example
For instance, prompting a 540B-parameter language model with just eight chain of thought exemplars achieves state of the art accuracy on the GSM8K benchmark of math word problems, surpassing even finetuned GPT-3 with a verifier. The before-and-after difference is stark on a single problem. A standard prompt gets a question like "Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now?" and jumps straight to an answer, sometimes wrong, with no visible working. A chain-of-thought prompt produces something closer to: "Roger started with 5 balls. 2 cans of 3 tennis balls each is 6 tennis balls. 5 + 6 = 11. The answer is 11," and gets the arithmetic right far more reliably, because each step is small enough for the model to get right on its own.
The zero-shot version of the same technique showed similarly large jumps without any worked examples at all: increasing the accuracy on MultiArith from 17.7% to 78.7% and GSM8K from 10.4% to 40.7% with large InstructGPT model (text-davinci-002), just from adding the instruction to think step by step before answering.
052 min
What this means for your work
For an engineer building a product feature on top of an LLM, chain-of-thought prompting is usually the first technique to try on any task that involves arithmetic, multi-step logic, or a decision with several conditions, before reaching for a more expensive fix like fine-tuning. The tradeoff a PM needs to understand is that the extra reasoning text is not free: every token the model writes as reasoning is a token of latency and cost paid before the user sees a final answer, so a feature that needs a fast, cheap response is a worse fit for the technique than a batch job with nobody waiting on it.
A founder evaluating whether a smaller, cheaper model is good enough for a task should know the technique is not a universal upgrade. Reasoning abilities emerge naturally in sufficiently large language models via a simple method called chain of thought prompting, which means applying it to a small model can produce longer, more confident-looking wrong answers rather than more correct ones, because the model is imitating the shape of reasoning without reliably doing the underlying steps correctly. A designer building a chat interface also has a real decision to make about whether to show the model's reasoning text to the user at all, since Turpin and colleagues in 2023 found that a model's stated reasoning can misrepresent what actually drove its answer, so a visible chain of thought reads as an explanation without necessarily being a reliable one.
A common extension worth knowing before reaching for fine-tuning is self-consistency: instead of generating one chain of reasoning, the system samples several independent chains for the same question at a higher temperature, then takes whichever final answer the majority of the chains agree on. It trades a proportional increase in cost, since every sampled chain is a separate generation, for a further accuracy gain on the same class of multi-step tasks, and it is a purely inference-time technique, so it needs no retraining and no new prompt examples beyond the ones chain-of-thought already uses.
061 min
What Chain-of-Thought Prompting costs
Chain-of-thought reasoning multiplies the number of tokens a model generates for a single answer, sometimes by five to ten times over a direct answer, and since most API pricing charges per output token, the technique directly multiplies the dollar cost and the latency of every call that uses it. On a task where a direct answer would have taken one second and cost a fraction of a cent, a chain-of-thought version that writes several sentences of reasoning first can take several seconds and cost several times as much, which matters at the scale of millions of requests even though it is negligible for a single call. Self-consistency's majority-vote variant multiplies that cost again, by however many chains are sampled, so it is usually reserved for tasks where the accuracy gain is worth paying for several independent generations per answer rather than one.
071 min
What Chain-of-Thought Prompting does not solve
Chain-of-thought prompting's biggest limitation is that its benefit only reliably appears in large models. We show how such reasoning abilities emerge naturally in sufficiently large language models via a simple method called chain of thought prompting, which is another way of saying the same technique tested on a small model often shows little benefit or none, because a small model can produce fluent-looking reasoning steps that do not actually track a correct solution path.
A separate, more troubling limit is that the reasoning a model shows is not guaranteed to be the real reason for its answer. CoT explanations can be heavily influenced by adding biasing features to model inputs--e.g., by reordering the multiple-choice options in a few-shot prompt to make the answer always "(A)"--which models systematically fail to mention in their explanations. A team that treats a model's chain of thought as a trustworthy audit trail, rather than as a plausible-sounding narrative the model also had to generate, can be misled about why a system actually produced a given answer.
081 min
Chain-of-Thought Prompting vs. nearby concepts
| Concept | What it names | How it differs |
|---|---|---|
| Not in the library yetChain-of-Thought Prompting | ||
| AIZero-Shot Learning | ||
| AIFew-Shot Learning | ||
| AIFine-Tuning |
The cleanest way to place chain-of-thought is by what it changes: it is a prompting strategy that can be layered onto either a zero-shot or a few-shot prompt, while fine-tuning is a different lever entirely, changing the model rather than the request sent to it.
091 min
How Chain-of-Thought Prompting changed since
Wei and colleagues introduced chain-of-thought prompting in January 2022, using worked examples that paired each question with reasoning text rather than just a final answer, and reported state-of-the-art results on math word problems using that approach. Kojima and colleagues showed four months later that the worked examples were not even necessary, achieving similarly large gains simply by adding the instruction to think step by step, with zero examples, a much cheaper technique to apply since it needs no curated examples at all.
By 2023, the technique's limits were becoming as well documented as its benefits. Turpin and colleagues' paper "Language Models Don't Always Say What They Think" showed that the reasoning a model displays is not a reliable window into the computation that actually produced its answer, shifting the field's framing of chain-of-thought from a transparency feature toward a performance technique whose explanations need independent verification.
?6 questions





