011 min
Few-Shot Learning at a glance
- What it is: Steering a model with examples written into the prompt, not by retraining it.
- Why it exists: Collecting labeled data and fine-tuning for every new task was too slow to be practical.
- What it costs: Every example eats context-window budget on every single request, forever.
- When it breaks: Accuracy can swing with example order or wording, not just example content.
- As of: October 2026, OpenAI is winding down fine-tuning access for new users, making prompting the default starting point.
021 min
The problem Few-Shot Learning solves
Before few-shot prompting, adapting a language model to a new task meant one of two slow routes. The first was collecting a labeled dataset large enough to fine-tune the model's weights on the new task, which could take days of data collection and a training run before a single result came back. The second was writing one natural-language instruction and hoping the model's zero-shot understanding of the task was good enough, with no way to show it what a correct answer actually looked like, and no real recourse when it guessed wrong in a consistent, predictable way.
GPT-3's 2020 paper framed the alternative directly: give the model a handful of worked examples inside the prompt itself, with no change to its weights at all. The paper's own abstract states the model is "applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model." That removed the days-long data-collection step for a huge range of tasks, at the cost of paying for those same examples again on every single request instead of once, during a training run that never has to happen at all.
032 min
How Few-Shot Learning works
Say a support team wants to automatically label incoming tickets as "billing," "technical," or "account." Zero-shot, the model gets only an instruction: classify this ticket into one of the three categories, using only its own general sense of what each word tends to mean. Few-shot, the prompt also includes two or three real tickets, each paired with its correct label, placed before the new ticket that actually needs classifying. The model's weights are never touched; it reads the examples fresh on every single request and infers the pattern from them in context, the way a person might infer a sorting rule from a few worked examples rather than from a dictionary definition of each category.
The GPT-3 paper draws the same three-way distinction formally: "(a) 'few-shot learning,' or in-context learning where we allow as many demonstrations as will fit into the model's context window (typically 10 to 100), (b) 'one-shot learning,' where we allow only one demonstration, and (c) 'zero-shot' learning, where no demonstrations are allowed and only an instruction in natural language is given to the model." The paper is explicit that nothing about the model itself changes between these three settings: "We emphasize that these 'learning' curves involve no gradient updates or fine-tuning, just increasing numbers of demonstrations given as conditioning."
The step where the interesting thing happens is inference time, not training time. A fine-tuned model has already baked the support-ticket pattern into its weights before it ever sees a real ticket, the way a worker who has been trained on hundreds of past tickets already has the categories internalized before their next shift starts. A few-shot-prompted model is doing something closer to pattern completion on the fly, using only the examples already sitting in its context as the signal for what "billing" versus "technical" should look like for this one conversation. That is also why the examples have to travel with every request: remove them, and the model has nothing left to infer the task from except the bare instruction, the same way the worker above would struggle on day one without ever having seen a single past ticket.
The examples also do something a bare instruction cannot: they pin down format as well as meaning. A zero-shot instruction can state the three category names perfectly clearly and still get inconsistent capitalization, punctuation, or explanation length back, because nothing in the prompt shows the model exactly what a finished answer should look like; a few-shot example shows both the category and the exact shape of a good response in one move.
041 min
A concrete example
The clearest illustrative contrast is a before-and-after on the same support-ticket task:
| Prompt style | What the model receives | Typical result |
|---|---|---|
| Zero-shot | Only the instruction and the new ticket | Reasonable guesses on obvious tickets, inconsistent category boundaries on ambiguous ones |
| Few-shot (3 examples) | The instruction, three labeled example tickets, then the new ticket | More consistent category boundaries, because the examples pin down exactly where "billing" ends and "account" begins |
That pattern shows up in GPT-3's own published numbers, not just as an illustration. On TriviaQA, a closed-book trivia-question benchmark, GPT-3 175B scored 64.3% in the zero-shot setting, 68.0% one-shot, and 71.2% few-shot, a result the paper calls "state-of-the-art relative to fine-tuned models operating in the same closed-book setting." The same paper's LAMBADA benchmark shows the jump is not always monotonic: zero-shot scored 76.2%, one-shot dipped slightly to 72.5%, and few-shot then jumped to 86.4%, an 18% relative gain over the previous state of the art. The one-shot dip is itself a real finding worth keeping, not an error: a single example can mislead a model about a task's exact format in a way that zero or several examples do not.
051 min
What this means for your work
A PM scoping a new AI feature gets a real decision out of this: a few-shot prompt can validate whether a task is even feasible for a model within a single afternoon, before anyone commits to collecting a fine-tuning dataset. That is why most teams prototype with prompting first and fine-tune only once the prompted version has already proven the task is worth solving, rather than committing to a training run on an unproven idea.
A designer building the prompt itself has to decide how many examples to include and in what order, and that choice is not cosmetic: because the GPT-3 paper's own numbers show few-shot results can swing based on which and how many examples appear, the examples chosen for a shipped feature should be representative of the actual range of real inputs, not just the two or three cleanest cases a designer happened to write first while drafting the prompt.
An engineer running evaluation has the sharpest version of this decision: a model's apparent accuracy on a few-shot task depends partly on the specific examples used to measure it, so an eval set built from only one ordering or one fixed set of examples will overstate how reliable the feature really is in production, where real inputs will not arrive in that same convenient, carefully-chosen order.
061 min
What Few-Shot Learning costs
Every example included in a few-shot prompt eats into the model's context window and is paid for again on every single request, which is the opposite of fine-tuning's cost shape: a one-time training cost followed by short, cheap prompts afterward. OpenAI's own documentation states the trade-off directly: fine-tuning lets a team "use shorter prompts with fewer examples and context data, which saves on token costs at scale and can be lower latency," compared with carrying the same examples in every few-shot call.
As of October 2026, that trade-off is shifting in prompting's favor for new projects: OpenAI's documentation states that the company "is winding down the fine-tuning platform," and that it "is no longer accessible to new users," though existing fine-tuned models remain available for inference. That makes few-shot prompting the only practical starting point for a new OpenAI-based feature today, whatever its long-run token cost, since the fine-tuning route is no longer open to new users at all.
071 min
What Few-Shot Learning does not solve
Few-shot prompting does not solve consistency, and the GPT-3 paper's own numbers make the point better than a warning could: on the LAMBADA benchmark, one-shot actually scored lower than zero-shot, 72.5% against 76.2%, before few-shot recovered to 86.4%. A single example, in other words, is not strictly safer than none; it can mislead the model about the exact shape of the task in a way that either no examples or several examples do not.
It also does not solve genuinely novel or highly specialized tasks, where no number of in-prompt examples substitutes for the kind of domain-specific training data a fine-tuning run could draw on. A well-known overclaim worth correcting directly: more examples does not reliably mean better results. The LAMBADA dip above is the concrete counter-example to that intuition, and it comes from the very paper most often cited to support it, so the overclaim is not even a secondhand misreading, it is a misreading of the paper's own numbers.
081 min
Few-Shot Learning vs. nearby concepts
The deciding fact between few-shot prompting and fine-tuning is whether the model's weights change. Few-shot prompting leaves the weights untouched and supplies examples at inference time, in the prompt; fine-tuning permanently updates the weights using a labeled dataset, after which the examples no longer need to travel with every request. Zero-shot learning is simply few-shot with the example count set to zero: same mechanism, same unchanged weights, just an instruction with nothing to pattern-match against.
Few-shot learning in computer vision is an older, different idea with a similar name but a different mechanism. Fei-Fei, Fergus, and Perona's 2006 paper "One-Shot Learning of Object Categories" showed a model could learn to recognize a new visual category from just one or a handful of labeled images, but it did so by updating the model's internal parameters on those examples through a learned prior, not by holding the examples in a text prompt at inference time. The LLM sense of few-shot learning, in-context learning via prompting, did not exist in anything like its current form until over a decade later.
092 min
A second case
A second, genuinely different case of few-shot learning predates LLM prompting by fourteen years: Fei-Fei, Fergus, and Perona's 2006 computer-vision paper, which opens with the problem directly: "Learning visual models of object categories notoriously requires hundreds or thousands of training examples. We show that it is possible to learn much information about a category from just one, or a handful, of images." Their system learned a brand-new object category from one or a few labeled images by updating a Bayesian prior built from categories it had already learned, so the model's actual internal parameters changed with each new example, unlike a GPT-3 prompt that leaves every weight untouched.
The variable that differs between the two cases is where the adaptation happens. GPT-3's few-shot learning adapts at inference time, inside a single prompt, with the underlying weights completely frozen; the 2006 vision system adapted its own internal model on the spot, with no separate "prompt" step and no frozen weights at all. What the contrast teaches is that "few-shot learning" names a goal, learning a new task or category from very few examples, not a specific mechanism for reaching it; LLM in-context learning and classical few-shot computer-vision learning both chase that same goal through almost entirely different machinery, fourteen years and one whole paradigm apart. The 2006 paper's title itself, naming the goal "one-shot learning of object categories," is a reminder that the one-shot and few-shot naming convention used for GPT-3 in 2020 was already established vocabulary in machine learning well before anyone applied it to a language model's prompt.
101 min
Where the evidence is contested
Few-shot prompting is less reliable than its benchmark numbers alone suggest, and this is a documented finding, not a hedge. Zhao, Wallace, Feng, Klein, and Singh (2021) showed that "the choice of prompt format, training examples, and even the order of the training examples can cause accuracy to vary from near chance to near state-of-the-art" on the same task with the same model. A separate study, Lu, Bartolo, Moore, Riedel, and Stenetorp (2021), isolated the ordering variable specifically and found that "the order in which the samples are provided can make the difference between near state-of-the-art and random guess performance," describing some orderings as simply "fantastic" and others as not.
Zhao and colleagues trace the instability to a specific, correctable bias: models lean toward predicting answers that sit "near the end of the prompt or are common in the pre-training data," and their proposed fix, recalibrating the model's output probabilities before using them, recovered "up to 30.0% absolute" in average accuracy. The practical stance while this is unsettled: a reported few-shot accuracy number is only as trustworthy as the example set and ordering it was measured with, and a team that evaluates a prompt on one fixed example order has measured one point in a range that the cited research shows can swing by 30 points or more.
111 min
How Few-Shot Learning changed since
"Few-shot learning" as a general machine-learning goal, learning a new task from very few examples, predates large language models by well over a decade. Fei-Fei, Fergus, and Perona's 2006 paper on one-shot object recognition is the clearest early landmark, aimed squarely at computer vision rather than language, and it worked by updating a model's own internal parameters on each new example rather than by holding examples inside any kind of text prompt.
The term carried over to language models with a different mechanism entirely once GPT-3 arrived in 2020. Brown and colleagues used "few-shot learning" for a model that reads examples straight from its own context window at inference time, with its weights completely frozen throughout, and that in-context sense is what almost everyone building AI features means by the term today. Current usage has mostly displaced the older vision sense in everyday conversation, even though the original 2006 meaning, learning from very few examples, is still technically accurate for both senses of the phrase.
122 min
Frequently asked questions
Is few-shot learning the same as fine-tuning? No. Few-shot learning puts examples in the prompt and leaves the model's weights untouched. Fine-tuning permanently updates the weights using a labeled dataset, after which the examples no longer need to be sent with every request.
When should I use few-shot instead of fine-tuning? Use few-shot prompting to validate whether a task is even feasible within hours, before committing to the dataset collection and training time fine-tuning requires. Move to fine-tuning once the prompted version works and the per-request example cost becomes the bottleneck.
What are the risks or limits of few-shot learning? Results can swing with the order and wording of the examples, not just their content, by as much as 30 percentage points in published research. A reported accuracy number is only as trustworthy as the specific example set it was measured with.
Does few-shot learning require GPUs or a fine-tune? No. Few-shot prompting runs on an unmodified, already-trained model; the examples are sent as plain text inside the prompt. No training run, GPU cluster, or labeled dataset collection is required to use it.
How much does few-shot learning cost to run? Every example in the prompt is billed as input tokens on every single request, so cost scales with how many examples are included times how often the feature is called, unlike fine-tuning's one-time training cost.
Is few-shot learning the same thing in computer vision and in language models? The goal is the same, learning from very few examples, but the mechanism differs. The 2006 computer-vision sense updates the model's own parameters per example; the language-model sense since 2020 keeps the model frozen and reads examples from the prompt instead.
Why does adding more examples sometimes make a model worse? A single misleading example can distort the model's sense of the task's exact format. GPT-3's own published results show one-shot scoring lower than zero-shot on one benchmark before few-shot recovered and surpassed both, so more examples is not a strict guarantee of better results.
?7 questions
Questions people ask
Is few-shot learning the same as fine-tuning?
When should I use few-shot instead of fine-tuning?
What are the risks or limits of few-shot learning?
Does few-shot learning require GPUs or a fine-tune?
How much does few-shot learning cost to run?
Is few-shot learning the same thing in computer vision and in language models?
Why does adding more examples sometimes make a model worse?
§6 sources
Sources
Brown, T. B., et al. (2020). Language Models are Few-Shot Learners. arXiv:2005.14165. (full text: )
Fei-Fei, L., Fergus, R., & Perona, P. (2006). One-Shot Learning of Object Categories. IEEE Transactions on Pattern Analysis and Machine Intelligence, 28(4).
Zhao, Z., Wallace, E., Feng, S., Klein, D., & Singh, S. (2021). Calibrate Before Use: Improving Few-Shot Performance of Language Models. arXiv:2102.09690.
Lu, Y., Bartolo, M., Moore, A., Riedel, S., & Stenetorp, P. (2021). Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity. arXiv:2104.08786.
Show all 6 sourcesShow fewer sources
OpenAI. Model Optimization Guide (prompting, fine-tuning, and their trade-offs).
Anthropic. Claude Prompting Best Practices (multishot / few-shot example guidance).





