Context Window

The total amount of text a language model can consider in a single request, measured in tokens and shared between the input and the response.

12 min read

Β· Also in

By Ravi SuranaUpdated 7 sources

Quick answer

~20 sec

A context window is the total amount of text a language model can consider at once, measured in tokens. It holds the instructions, the conversation so far, any documents or tool results supplied, and the response being generated. When the total exceeds the limit, the request fails or older content has to be removed.

012 min

How a context window works

Take one running example: an assistant that answers questions about a company's refund policy.

The application sends a request containing three things. A system prompt setting out how the assistant should behave, say 600 tokens. The refund policy document, 4,000 tokens. And the customer's question, 20 tokens. A token is a common chunk of text, usually a few characters, so a document of roughly 3,000 English words lands near 4,000 tokens.

All 4,620 tokens are laid out as one sequence. The model computes over that whole sequence at once, and every token can draw on every other token. Then it produces an answer, one token at a time, with each new token appended to the same sequence. If the answer runs to 300 tokens, the final sequence is 4,920 tokens long.

Why there is a limit at all

Attention compares every token against every other token. Vaswani and colleagues recorded the consequence in the 2017 Transformer paper: the work in an attention layer grows with the square of the sequence length. Dao and colleagues restate it in the FlashAttention paper as both time and memory being quadratic in sequence length. Doubling the sequence roughly quadruples the attention arithmetic and the memory it needs.

The published limit for a model is therefore a decision about cost and hardware, not a property of the architecture. Nothing in the design forbids a longer sequence. Serving it economically at a given latency is the constraint.

Where the customer's second question goes

The model keeps nothing between requests. When the customer asks a follow-up, the application re-sends the system prompt, the policy document, the first question, the first answer, and the new question. The window is now carrying 5,000 tokens of history that did not exist one turn ago.

That is the property that surprises teams most often. A conversation does not accumulate inside the model. It accumulates inside every request your application makes, and each turn pays for all of it again.

022 min

What actually counts against the window

Anthropic's developer documentation is explicit about this for its own models, and the list is longer than most teams assume. As of September 2026 it states that everything in the request counts: the system prompt, every message including tool results, images and documents, and the tool definitions themselves. The output counts too, including any extended thinking the model produces.

Two entries on that list are the ones that catch people.

Tool definitions. An agent with thirty tools carries the full text of thirty tool schemas in every single request, whether or not it calls any of them. That cost is paid before the user has typed anything.

Tool results. A search tool that returns ten results of 800 tokens each adds 8,000 tokens to the window, and those results stay in the history for the rest of the session unless the application removes them.

The documented limits, again as of September 2026, are 1 million tokens for the current Claude models and 200,000 for several earlier ones. If the input alone exceeds the window, the API returns a 400 error reading "prompt is too long". Anthropic documents two mechanisms for staying under it: server-side compaction, which summarises earlier parts of a conversation, and context editing, which clears old tool results.

These figures describe Anthropic's own product and nothing else. Every provider publishes different limits and counts tokens slightly differently, so the number to design against is the one in your provider's documentation on the day you design.

032 min

What a large context window does not solve

A model that accepts 200,000 tokens is not a model that uses 200,000 tokens equally well. Two pieces of published evidence say so directly.

Liu and colleagues (2023) tested multi-document question answering and key-value retrieval while moving the relevant passage around the input. Performance was highest when the needed information sat at the beginning or the end of the context, and dropped significantly when it sat in the middle. They named the finding lost in the middle, and reported it even for models built specifically for long contexts.

Hsieh and colleagues (2024) built RULER, a benchmark that goes past the simple test of hiding one fact in a long document. Every model they tested claimed a context size of 32,000 tokens or more. Only half maintained satisfactory performance at 32,000. Almost all scored near-perfectly on the simple hidden-fact test and then fell away sharply as real task length grew.

Anthropic's own documentation names the same pattern for its models, calling it context rot: as token count grows, accuracy and recall degrade. The advice it draws from this is worth repeating, because it inverts the instinct. More context is not automatically better, and curating what goes in matters as much as how much room there is.

So the failure mode is quiet. Nothing errors. The model answers, the answer is wrong in a way that looks plausible, and the reason is that the one paragraph that mattered was in position 40,000 of 90,000.

042 min

What the context window means for your work

For an engineer: evaluate at the length you actually ship. A retrieval system tested with a 2,000-token context and deployed with a 60,000-token one has not been tested. The RULER results say the gap between those two conditions is where models separate. Build the evaluation set at production length, and include cases where the answer sits in the middle of the input.

For an engineer, second decision: order the input deliberately. Given the lost-in-the-middle finding, the beginning and the end of the context are the strong positions. Retrieved passages that matter most belong at one of those ends rather than in the order the search index returned them.

For a product manager: long conversations degrade before they fail. The user-visible symptom is not an error message. It is an assistant that forgets a constraint the user set twenty turns ago, gets slower, and costs more per turn. If your product has long sessions, the question to put on the roadmap is what gets summarised or dropped, and who decides.

For a founder: a larger window is not a substitute for retrieval. Sending an entire knowledge base on every request is simpler to build and worse on three axes at once: cost, latency and accuracy at depth. Retrieval exists to put a small number of relevant passages in the strong positions of a short context. A bigger window changes how much you can send, not how much you should.

051 min

What a context window costs

Three separate costs move together when the window fills, and teams usually notice only the first.

Money. Input tokens are billed on every request. A 50,000-token prompt re-sent across a 30-turn conversation is 1.5 million input tokens, even though the user typed a few hundred words. Prompt caching changes the price of repeated tokens but not whether they occupy the window; Anthropic's documentation states that cached prefixes still count toward the limit.

Latency. Time to first token rises with input length, because the whole input is processed before the first output token appears. Long prompts make an interface feel slow at exactly the moment the user is waiting for something to happen.

Accuracy. This is the one that has no line on the invoice. The RULER results and the lost-in-the-middle results both point the same way: quality falls with length before the limit is reached, and it falls silently.

The practical consequence is that context should be treated as a budget rather than a container. Something that is in the window is spending all three currencies, every turn, until it is removed.

062 min

Context window vs. nearby concepts

Compared withThe one fact that decides
AITraining dataTraining data shaped the weights once and is not retrievable as text. The context window is text supplied at request time. A model can quote a document in its context exactly; it cannot reliably quote its training data at all.
AIRetrieval-Augmented GenerationNot alternatives. Retrieval is a method for deciding what goes into the context window. The window is the space retrieval is selecting for. A larger window makes retrieval less urgent, never unnecessary.
Not in the library yetRate limitA rate limit caps tokens per minute across requests. The context window caps tokens within one request. You can hit either without touching the other.
Not in the library yetOutput limitMost providers cap output tokens separately and more tightly than input. A 1-million-token window does not mean a 1-million-token answer.
PsychologyWorking memoryA human analogy that misleads on the important point. Human working memory holds roughly four items and degrades by forgetting the oldest. A context window holds hundreds of thousands of tokens and degrades by weighting the middle less, which is a different failure with different fixes.

When somebody asks whether to use a long context or retrieval, the deciding fact is how much of the corpus could possibly be relevant to one question. If it is a few passages, retrieve them. If the whole document genuinely has to be read end to end, send it and accept the cost.

072 min

A second case: an agent loop

The refund-policy assistant fills its window with text a person chose to send. An agent fills it with text nobody chose.

Consider a coding agent working through a task. It starts with a system prompt and a set of tool definitions, perhaps 5,000 tokens. It reads a file: 3,000 tokens of tool result. It searches the repository: 2,000 tokens of results, most of them irrelevant. It reads two more files, runs a test suite and receives 4,000 tokens of output, most of it a stack trace it needed one line from. Twenty steps in, the window holds 80,000 tokens, and the great majority is tool output that mattered for one step and has been carried ever since.

One variable separates this from the first case: who decides what enters the window. In a chat application a human types, and volume is bounded by typing speed. In an agent loop the model's own tool calls generate the input, and volume is bounded only by how much the tools return.

What the contrast teaches is not visible in either case alone. Context management in a chat product is a compression problem, and summarising older turns solves most of it. Context management in an agent is a selection problem: the right fix is usually to make tools return less, or to clear results once they have been used, rather than to summarise after the fact. Anthropic documents tool-result clearing as a distinct mechanism from compaction for this reason.

082 min

Where the evidence is contested

The disagreement here is between vendor measurements and independent benchmarks, and it is unusually clean because both were published in the same year.

The Gemini 1.5 technical report from Google (2024) claims near-perfect retrieval, above 99 percent, up to at least 10 million tokens, and presents this as a generational step over what it names as the then-current alternatives: Claude 3.0 at 200,000 tokens and GPT-4 Turbo at 128,000.

RULER, from Hsieh and colleagues at NVIDIA in the same year, reports that models score nearly perfectly on exactly the kind of simple retrieval test that claim rests on, while degrading substantially on multi-hop tracing and aggregation at the same lengths.

Stated fairly, neither result is wrong. They measure different things. Finding one distinctive sentence in 10 million tokens is a retrieval task. Tracking several entities across a long document, or aggregating facts scattered through it, is a reasoning task performed over a long input. A model can be excellent at the first and mediocre at the second, and the published numbers say that many are.

Where this leaves a practitioner is specific rather than vague. Treat an advertised context length as a ceiling on what will be accepted, not a promise about quality, and discount any long-context claim whose evidence is a single-fact retrieval test. The only number that settles it for your product is one you measure on your own task at your own length.

091 min

How the context window changed since

Published limits have moved by roughly four orders of magnitude in under a decade, and what changed along the way is more interesting than the numbers.

The original Transformer paper in 2017 was not framed in terms of a context window at all. It trained on sentence pairs, and sequence length was a training detail rather than a product specification.

By 2024 the length had become a headline feature. Google's Gemini 1.5 report names the competitive field precisely: Claude 3.0 at 200,000 tokens and GPT-4 Turbo at 128,000, against its own claim of usable context into the millions.

As of September 2026, Anthropic documents 1 million tokens as the default window on its current models, with no special header required. The frontier of the specification has moved from hundreds of tokens to millions.

What has not moved in step is quality across that length. The 2023 lost-in-the-middle result and the 2024 RULER results both landed while the advertised numbers were climbing fastest, and both said the same thing: capacity grew faster than reliable use of that capacity. Vendor documentation now says so too, which is the real change in usage. The question asked about a model in 2023 was how much context it accepts. The question asked in 2026 is how much of it the model actually uses well.

?8 questions

Questions people ask

What counts toward the context window?

Everything in the request and the response. Anthropic's documentation lists the system prompt, every message, tool results, images, documents, the tool definitions themselves, and the generated output including extended thinking.

How many words is a token?

Roughly three quarters of an English word on average, so 1,000 tokens is around 750 words. The ratio varies by language and is worse for code, numbers and languages that do not use spaces between words.

Is a bigger context window always better?

No. Anthropic's own documentation states that accuracy and recall degrade as token count grows. RULER (2024) found that only half the models tested held up at 32,000 tokens despite all of them claiming that length or more.

Do I still need RAG if the context window is large?

Usually yes. Retrieval decides what enters the window, which controls cost, latency and accuracy at depth. A larger window changes how much you can send, not how much is worth sending.

What happens when the context window is exceeded?

Anthropic's API returns a 400 error reading "prompt is too long" when the input alone is over the limit. Chat interfaces often drop the oldest turns instead, which is why long chats appear to forget their own beginnings.

Why do long conversations get more expensive?

The model keeps no state between requests, so the whole conversation is re-sent every turn and billed as input each time. A 50,000-token history across 30 turns is 1.5 million input tokens.

Where should important information go in a long prompt?

At the beginning or the end. Liu and colleagues (2023) found performance is highest when relevant information sits at either edge of the context and drops significantly when it sits in the middle.

Is the context window the same as the model's memory?

No. It is working space for one request, not storage. Nothing in the window survives to the next request unless your application sends it again, and nothing in it changes the model's weights.

Β§7 sources

Sources

  1. Vaswani, A. et al. (2017). Attention Is All You Need. arXiv:1706.03762

  2. Liu, N. F. et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172

  3. Hsieh, C.-P. et al. (2024). RULER: What's the Real Context Size of Your Long-Context Language Models? arXiv:2404.06654

  4. Gemini Team, Google (2024). Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv:2403.05530

Show all 7 sources
  1. Dao, T. et al. (2022). FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. arXiv:2205.14135

  2. Anthropic. Context windows (developer documentation, accessed September 2026)

  3. Anthropic. Effective context engineering for AI agents (accessed September 2026)

Keep reading

More from AI

All of AI
All of AI