Chain of Thought

A prompting technique that asks a large language model to write out its intermediate reasoning steps before giving a final answer, instead of jumping straight to the answer.

9 min read

By Ravi SuranaUpdated

Quick answer

~20 sec

Chain-of-thought prompting asks a large language model to write out its intermediate reasoning steps before giving a final answer, instead of jumping straight to the answer. Adding a few worked examples, or even the phrase "let's think step by step," measurably improves accuracy on tasks that require multiple steps, like arithmetic and logic word problems.

011 min

Key takeaways

  • What it is: prompting a model to show its reasoning steps before its final answer
  • Why it exists: a direct answer skips steps a model needs to get multi-step problems right
  • What it costs: more output tokens and higher latency for every answer
  • When it breaks: the gains largely appear only in sufficiently large models
  • As of: the original 2022 papers reported the largest gains on math and logic benchmarks

021 min

The problem Chain-of-Thought Prompting solves

Before chain-of-thought prompting, the standard way to get a large language model to answer a question was to show it a few examples of a question paired directly with its final answer, with nothing in between, and then ask a new question the same way. That worked well for a task a model could answer in a single step, like classifying a sentence's sentiment, but it failed on a task that needs several steps of reasoning chained together, like a multi-step arithmetic word problem. Asked directly for the final number, a model would frequently skip a step or combine two numbers incorrectly, the same way a person asked to do multi-digit arithmetic entirely in their head, with no scratch paper, makes more errors than the same person working it out on paper. The model had no equivalent of scratch paper: standard prompting gave it no space to work anything out before it had to commit to a final answer. Simply making the model bigger did not fix this on its own; a larger model prompted the same direct way still skipped the same intermediate steps, which suggested the problem was in how the question was asked, not only in how much the model knew.

032 min

How Chain-of-Thought Prompting works

Chain-of-thought prompting changes what the examples in a prompt look like, not the model itself. Instead of pairing each example question directly with its final answer, the prompt pairs each example question with a short written explanation of the reasoning that leads to the answer, and only then the answer itself. When the model is later given a new question in the same format, it continues the pattern: writing out its own reasoning steps before producing a final answer.

The mechanism works because a large language model, almost always built on the Transformer architecture, generates its response one token at a time, each new token conditioned on every token that came before it, including the tokens it has just generated itself. Writing an intermediate reasoning step first means every later token, including the final answer, is generated while conditioned on that reasoning text sitting in front of it, functioning like a working memory the model can refer back to. Skipping straight to the answer removes that scratch space entirely: the model has to arrive at the right answer using only its internal computation for a single answer token, with no intermediate text to condition on. A version of the same idea needs no worked examples at all. Kojima and colleagues showed that simply adding the words "let's think step by step" before a model's answer, with no example reasoning shown at all, triggers the same behaviour, because instructive language alone is often enough to make the model produce reasoning text instead of jumping to an answer.

041 min

A concrete example

For instance, prompting a 540B-parameter language model with just eight chain of thought exemplars achieves state of the art accuracy on the GSM8K benchmark of math word problems, surpassing even finetuned GPT-3 with a verifier. The before-and-after difference is stark on a single problem. A standard prompt gets a question like "Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now?" and jumps straight to an answer, sometimes wrong, with no visible working. A chain-of-thought prompt produces something closer to: "Roger started with 5 balls. 2 cans of 3 tennis balls each is 6 tennis balls. 5 + 6 = 11. The answer is 11," and gets the arithmetic right far more reliably, because each step is small enough for the model to get right on its own.

The zero-shot version of the same technique showed similarly large jumps without any worked examples at all: increasing the accuracy on MultiArith from 17.7% to 78.7% and GSM8K from 10.4% to 40.7% with large InstructGPT model (text-davinci-002), just from adding the instruction to think step by step before answering.

052 min

What this means for your work

For an engineer building a product feature on top of an LLM, chain-of-thought prompting is usually the first technique to try on any task that involves arithmetic, multi-step logic, or a decision with several conditions, before reaching for a more expensive fix like fine-tuning. The tradeoff a PM needs to understand is that the extra reasoning text is not free: every token the model writes as reasoning is a token of latency and cost paid before the user sees a final answer, so a feature that needs a fast, cheap response is a worse fit for the technique than a batch job with nobody waiting on it.

A founder evaluating whether a smaller, cheaper model is good enough for a task should know the technique is not a universal upgrade. Reasoning abilities emerge naturally in sufficiently large language models via a simple method called chain of thought prompting, which means applying it to a small model can produce longer, more confident-looking wrong answers rather than more correct ones, because the model is imitating the shape of reasoning without reliably doing the underlying steps correctly. A designer building a chat interface also has a real decision to make about whether to show the model's reasoning text to the user at all, since Turpin and colleagues in 2023 found that a model's stated reasoning can misrepresent what actually drove its answer, so a visible chain of thought reads as an explanation without necessarily being a reliable one.

A common extension worth knowing before reaching for fine-tuning is self-consistency: instead of generating one chain of reasoning, the system samples several independent chains for the same question at a higher temperature, then takes whichever final answer the majority of the chains agree on. It trades a proportional increase in cost, since every sampled chain is a separate generation, for a further accuracy gain on the same class of multi-step tasks, and it is a purely inference-time technique, so it needs no retraining and no new prompt examples beyond the ones chain-of-thought already uses.

061 min

What Chain-of-Thought Prompting costs

Chain-of-thought reasoning multiplies the number of tokens a model generates for a single answer, sometimes by five to ten times over a direct answer, and since most API pricing charges per output token, the technique directly multiplies the dollar cost and the latency of every call that uses it. On a task where a direct answer would have taken one second and cost a fraction of a cent, a chain-of-thought version that writes several sentences of reasoning first can take several seconds and cost several times as much, which matters at the scale of millions of requests even though it is negligible for a single call. Self-consistency's majority-vote variant multiplies that cost again, by however many chains are sampled, so it is usually reserved for tasks where the accuracy gain is worth paying for several independent generations per answer rather than one.

071 min

What Chain-of-Thought Prompting does not solve

Chain-of-thought prompting's biggest limitation is that its benefit only reliably appears in large models. We show how such reasoning abilities emerge naturally in sufficiently large language models via a simple method called chain of thought prompting, which is another way of saying the same technique tested on a small model often shows little benefit or none, because a small model can produce fluent-looking reasoning steps that do not actually track a correct solution path.

A separate, more troubling limit is that the reasoning a model shows is not guaranteed to be the real reason for its answer. CoT explanations can be heavily influenced by adding biasing features to model inputs--e.g., by reordering the multiple-choice options in a few-shot prompt to make the answer always "(A)"--which models systematically fail to mention in their explanations. A team that treats a model's chain of thought as a trustworthy audit trail, rather than as a plausible-sounding narrative the model also had to generate, can be misled about why a system actually produced a given answer.

081 min

Chain-of-Thought Prompting vs. nearby concepts

ConceptWhat it namesHow it differs
Not in the library yetChain-of-Thought PromptingA prompting technique, not a property of the model or the data it sawWhat it names: Asking a model to show reasoning steps before its final answer
AIZero-Shot LearningChain-of-thought can be combined with zero-shot prompting, as in "let's think step by step" with no examples at allWhat it names: A model performing a task it saw no worked examples of in the prompt
AIFew-Shot LearningChain-of-thought few-shot prompting is the specific case where those worked examples include reasoning, not just answersWhat it names: A model performing a task after seeing a small number of worked examples in the prompt
AIFine-TuningChanges the model itself; chain-of-thought changes only the prompt, and needs no retrainingWhat it names: Adjusting a model's own weights on new training examples

The cleanest way to place chain-of-thought is by what it changes: it is a prompting strategy that can be layered onto either a zero-shot or a few-shot prompt, while fine-tuning is a different lever entirely, changing the model rather than the request sent to it.

091 min

How Chain-of-Thought Prompting changed since

Wei and colleagues introduced chain-of-thought prompting in January 2022, using worked examples that paired each question with reasoning text rather than just a final answer, and reported state-of-the-art results on math word problems using that approach. Kojima and colleagues showed four months later that the worked examples were not even necessary, achieving similarly large gains simply by adding the instruction to think step by step, with zero examples, a much cheaper technique to apply since it needs no curated examples at all.

By 2023, the technique's limits were becoming as well documented as its benefits. Turpin and colleagues' paper "Language Models Don't Always Say What They Think" showed that the reasoning a model displays is not a reliable window into the computation that actually produced its answer, shifting the field's framing of chain-of-thought from a transparency feature toward a performance technique whose explanations need independent verification.

?6 questions

Questions people ask

Is chain-of-thought prompting the same as fine-tuning?

No. Chain-of-thought only changes the prompt sent to a model; fine-tuning changes the model's own weights and requires a training run, not just a different request.

When should I use chain-of-thought instead of a direct prompt?

On tasks with multiple reasoning steps, like arithmetic or multi-condition logic, where a direct answer tends to skip a step. For single-step tasks it adds cost with little benefit.

What are the risks or limits of chain-of-thought prompting?

It reliably helps only in sufficiently large models, and the reasoning a model shows is not guaranteed to reflect what actually produced its answer, so it should not be treated as a trustworthy explanation.

Does chain-of-thought require fine-tuning or a vector database?

No. It works entirely through the prompt, either with a few worked examples that include reasoning or, in the zero-shot version, a single instruction like "let's think step by step."

How much does chain-of-thought prompting cost to run?

It multiplies output tokens, sometimes five to ten times over a direct answer, so cost and latency scale directly with how much reasoning text the model generates before its final answer.

Is chain-of-thought the same as zero-shot reasoning?

They are related but distinct: the original chain-of-thought technique used worked examples with reasoning, while zero-shot chain-of-thought triggers similar behaviour with a single instruction and no examples at all.

Keep reading

More from AI

All of AI
All of AI