Zero-Shot Learning

A model correctly handling a class or task it saw zero labeled training examples of, by transferring from a description or related task.

12 min read

By Ravi SuranaUpdated 3 sources

Quick answer

~20 sec

Zero-shot learning is when a model correctly handles a class or task it saw zero labeled examples of, transferring instead from a description, an attribute list, or a related task. A photo classifier trained only on common animals can still name a narwhal in a new photo, given just the word narwhal, because it compares meaning, not memorized pixels.

011 min

The problem Zero-Shot Learning solves

Before zero-shot learning, an image or text classifier could only recognize the classes it had been shown during training. A photo-tagging system built to label ten thousand product categories needed labeled photos for all ten thousand before it shipped, and adding an eleven-thousandth category meant collecting, labeling and retraining rather than typing a new word into a menu. The label set was fixed at training time, and nothing could be added to it without a new labeled dataset and a new training run.

This naive approach, supervised learning with a fixed label set, fails as soon as the world produces a class faster than anyone can label it. Wildlife-monitoring cameras encounter species nobody staged a photo shoot for. Customer-support systems meet new product lines the week they launch. Content-moderation systems meet new slang before a labeling team can catch up. Few-shot learning eased the pain by needing only a handful of labeled examples per new class instead of thousands, but it still needs at least one labeled example for every class it must recognize. Zero-shot learning removes that floor: the system needs the class described, not shown.

021 min

Where Zero-Shot Learning came from

Hugo Larochelle, Dumitru Erhan and Yoshua Bengio gave the problem its first formal treatment in 2008, at the University of Montreal, under the name zero-data learning. Their AAAI paper trained a character-recognition model to predict classes it had never seen labeled examples of, using attribute vectors that described each class instead (Larochelle, Erhan and Bengio, 2008).

A year later, Mark Palatucci, Dean Pomerleau, Geoffrey Hinton and Tom Mitchell published Zero-Shot Learning with Semantic Output Codes at NeurIPS 2009, and that title is the name that stuck. Their paper explicitly builds on Larochelle and colleagues' work on zero-data learning and introduces the semantic output code classifier: a knowledge base of semantic properties, drawn from outside the training set, stands in for labeled examples of the new class. Their case study decoded which word a person was thinking about from fMRI brain images, for words the classifier had never been trained to decode (Palatucci, Pomerleau, Hinton and Mitchell, 2009).

032 min

How Zero-Shot Learning works

Zero-shot learning works by giving every class two representations instead of one: the input itself, and a description of the class that does not depend on having seen an example of it. Both representations are mapped into the same embedding, a list of numbers that stands for meaning, positioned so that similar things sit close together and different things sit far apart. Once inputs and class descriptions share that space, recognizing a new class becomes a comparison against nearby points, not a lookup in a table of labels the model was trained on.

Take a photo-tagging system trained on labeled pictures of common animals: dogs, cats, horses, sparrows. A user uploads a photo of a narwhal, a species with zero labeled photos anywhere in the training set. A conventional classifier has no slot for narwhal and can only guess among the classes it knows. A zero-shot system instead takes the image, runs it through a vision encoder trained during the labeled phase, and produces an embedding for that photo. Separately, it takes the word narwhal, or a short description such as a whale with a single long spiral tusk, and runs it through a matching encoder for text, producing an embedding for the description.

The step where the interesting thing happens

The system compares the photo embedding against the description embedding, along with the embeddings of every other candidate class name, and picks whichever description sits closest in that shared space. Nothing about narwhal was in the training labels. What was in the training data, indirectly, was enough shared structure between images and language, or between classes and their attributes, for the encoders to learn to place a photo of a tusked whale near the words that describe one, because they had already learned to place thousands of other photos near their own descriptions during training.

This is why the framing matters more than the plumbing. Larochelle, Erhan and Bengio's attribute vectors, Palatucci and colleagues' semantic output codes, and a modern image-text embedding are three different implementations of the same idea: replace "have I seen a labeled example of this class" with "can I compute how close this input sits to a description of this class." The side information, whether attributes, semantic codes, or a text embedding, is what does the work that labeled examples would otherwise have to do.

042 min

A concrete example

OpenAI's CLIP, described by Radford and colleagues in 2021, is zero-shot learning running at a much larger scale than either of the founding papers imagined. CLIP trains a single system on 400 million image-caption pairs collected from the internet, so that a photo and its caption end up close together in a shared embedding space, the same idea as the narwhal example above but learned from captions instead of hand-built attribute vectors. Both of CLIP's encoders are built from a transformer, the architecture most modern text and image embedding models use to turn a sequence into a single vector.

To recognize a new image class after training, CLIP never sees a single labeled example of it. It is given only the class name, embeds that name as if it were a caption, and picks whichever class name's embedding sits closest to the photo's embedding. Radford and colleagues report that the best CLIP model "improves accuracy on ImageNet from a proof of concept 11.5% to 76.2% and matches the performance of the original ResNet-50 despite using none of the 1.28 million crowd-labeled training examples available for this dataset" (Radford et al., 2021). ResNet-50 needed all 1.28 million of those labeled photos to reach a comparable score. CLIP needed zero of them, because the class names ImageNet uses were enough to place each photo correctly in the space CLIP had already learned from captions.

The gap this closes is real but not total. CLIP's zero-shot number is a match for one specific supervised baseline on one dataset, not a claim that zero-shot always equals supervised performance. On many of the more than thirty datasets Radford and colleagues test, CLIP's zero-shot accuracy trails a model trained directly on that dataset's own labels.

052 min

A second case

CLIP shows zero-shot learning in computer vision, where the side information is a natural-language caption. Palatucci, Pomerleau, Hinton and Mitchell's 2009 case study shows the same idea in a domain with no images and no captions at all: decoding which word a person is thinking about from an fMRI scan of their brain activity.

Their semantic output code classifier does not predict a word directly. It first predicts a vector of semantic features for the word, such as how strongly the concept relates to being an animal, being edible, or fitting in a hand, using feature norms collected by asking people to describe common nouns in a large, separate survey. Words the fMRI classifier had never been trained to decode still have a position in that same feature space, because the feature norms cover the word whether or not any brain-imaging data for it exists. The classifier then reports whichever untrained word's feature vector is closest to the one it just predicted from the scan.

The variable that changes between the two cases is the source of the side information. CLIP inherits its shared space from hundreds of millions of image-caption pairs scraped from the web. Palatucci and colleagues build theirs from a small, purpose-collected survey of human semantic judgments, decades before large web-scraped datasets were an option. What the two cases teach together is that zero-shot learning does not depend on a particular kind of side information, or a particular data source. It depends only on the side information and the input sharing a space in which distance carries meaning.

062 min

What this means for your work

A PM scoping a content-moderation feature has to decide how the system will handle a policy category invented next quarter, one nobody can label examples of yet. Betting entirely on supervised classifiers means every new category is a data-collection project before it is a shipped feature. Scoping the system so its classifier accepts a text description of a new category as input, the same shared-embedding approach CLIP uses, turns that into a config change instead of a retraining cycle, at the cost of lower accuracy than a category the team actually had time to label examples for.

An engineer building a product-search feature faces the mirror version of the same decision. A fixed label set is simpler to test and simpler to explain when it fails, but it breaks the day the catalog adds a category the model never saw named during training. A zero-shot layer, or a hybrid that falls back to zero-shot only for categories below some volume threshold, costs an extra encoder pass per query and adds a failure mode that is harder to debug: a wrong answer traceable to two vectors being closer than they should have been, not to a missing label.

A founder deciding where to spend the first labeling budget can use zero-shot accuracy as the signal for where hand-labeled data will pay for itself. Where a zero-shot model already reaches an acceptable score, as CLIP does on general object categories, added labeled data buys little. Where zero-shot lags badly, usually on categories that need domain vocabulary a general model was never trained on, such as recognizing a specific defect type on a manufacturing line, that is where labeling spend belongs first.

072 min

What Zero-Shot Learning costs

Zero-shot learning trades labeling cost for compute cost, and the trade is easy to miss because the labeling cost was visible on a project plan and the compute cost is buried in a bill later. At inference, a zero-shot system runs an extra encoder pass for the input and, in most implementations, one embedding lookup per candidate class name rather than a single forward pass through a fixed classification head. For a system choosing among a few dozen classes, that difference is negligible. For a system choosing among tens of thousands of candidate descriptions, comparing against every one of them at query time adds real latency, which is why production systems narrow the candidate set first with a cheaper method and use the zero-shot comparison only to rank what is left. Feeding a long list of candidate class descriptions straight into a language model's prompt instead runs into its context window, the fixed budget of text a model can attend to at once, which is another reason large candidate sets go through a separate embedding comparison rather than raw prompt text.

The other cost is upfront and one-time: building the shared embedding space at all. CLIP's version needed 400 million image-caption pairs. Palatucci and colleagues needed a purpose-run survey of human semantic judgments instead. Neither of those costs falls on the team using the resulting pretrained model, but they explain why most teams building a zero-shot feature reuse an existing encoder rather than building one from scratch.

082 min

What Zero-Shot Learning does not solve

Zero-shot learning fails first on fine-grained distinctions the side information cannot capture. Telling a photo of a crow from a photo of a raven from a class name alone is hard, because the words crow and raven sit close together in most text embeddings and the actual visual difference, mostly size and tail shape, is not written into either name. A classifier asked to make that call zero-shot will guess close to chance, while a model with even a handful of labeled examples of each bird usually will not.

It also fails when the side information itself is wrong or missing. A model built on attribute vectors written by hand inherits every gap in that attribute list. A model built on a text encoder trained mostly on English inherits its blind spots for classes whose names or descriptions rarely appear in that training text. Neither failure looks like an error message. Both look like a confident wrong answer, because the system still returns whichever candidate scored highest, even when every candidate scored badly.

The overclaim worth naming directly: zero-shot learning does not mean the system needed no data at all. It needed a large pretraining phase, on captions or attribute norms or a related task, and it needed that phase to actually cover something close to the new class. Zero refers to labeled examples of the specific class at hand, not to data in general.

092 min

Zero-Shot Learning vs. nearby concepts

The deciding fact between zero-shot, few-shot and ordinary supervised learning is how many labeled examples of the target class the system sees before it has to answer.

Labeled examples of the target classHow it generalizes
AISupervised learningLearns patterns directly from labeled examples of that exact classLabeled examples of the target class: Many, collected and labeled before training
AIFew-shot learningCombines the few examples with a pretrained representationLabeled examples of the target class: A handful, given at inference or fine-tuning time
Not in the library yetZero-shot learningTransfers from a description, an attribute list, or a related taskLabeled examples of the target class: None

Zero-shot prompting a chat model, meaning asking it a question with no worked examples in the prompt, is often described as a separate trick. It is not. It is the same underlying idea, applied to a large language model that was pretrained on enough text to already encode a description-like representation of most classes and tasks a user might ask about. Few-shot prompting is the prompting-world name for the same relationship few-shot learning has to zero-shot learning: adding a couple of worked examples to the prompt narrows the model down the way a few labeled examples narrow down any few-shot classifier.

Retrieval-augmented generation solves an adjacent but different problem. Instead of transferring from a description already baked into the model's weights during pretraining, it looks up fresh documents at query time and hands them to the model as extra context. The two combine well: retrieve a description of an unfamiliar class or topic, then let the model reason over that description zero-shot.

?5 questions

Questions people ask

Is zero-shot learning the same as few-shot learning?

No. Few-shot learning still needs a handful of labeled examples of the new class before it answers. Zero-shot learning needs none, relying instead on a description, an attribute list, or a related task the model already learned.

Is zero-shot prompting a different concept from zero-shot learning?

No, it is the same idea applied to a prompted language model. The model transfers from what it learned during pretraining to a task it was never explicitly trained on, using only the instructions in the prompt.

Does zero-shot learning need any training data at all?

Yes. It needs a large pretraining phase on captions, attribute norms, or a related task, so the model already has a shared space to compare against. Zero refers only to labeled examples of the specific new class.

How accurate is zero-shot learning compared to a trained classifier?

It varies by task. CLIP's zero-shot ImageNet accuracy matched a supervised ResNet-50 in 2021, but on many other benchmarks in the same paper, zero-shot trailed a model trained directly on that benchmark's own labeled data.

When should I collect labeled data instead of using zero-shot learning?

When the classes are fine-grained enough that their names or descriptions do not capture the real difference, such as visually similar species, or when zero-shot accuracy on your specific categories tests out too low to use.

Β§3 sources

Sources

  1. Larochelle, H., Erhan, D. and Bengio, Y. (2008). Zero-data Learning of New Tasks. AAAI Conference on Artificial Intelligence

  2. Palatucci, M., Pomerleau, D., Hinton, G. E. and Mitchell, T. M. (2009). Zero-Shot Learning with Semantic Output Codes. Advances in Neural Information Processing Systems 22 (NeurIPS 2009)

  3. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G. and Sutskever, I. (2021). Learning Transferable Visual Models From Natural Language Supervision (CLIP). arXiv:2103.00020

Keep reading

More from AI

All of AI
All of AI