Active Learning

A machine learning approach where the model selects which unlabeled examples to have labeled next, usually the ones it is least certain about, to reach higher accuracy with fewer labels.

8 min read

Β· Also in

By Ravi SuranaUpdated 1 source

Quick answer

~20 sec

Active learning is a machine learning approach where the model itself picks which unlabeled examples most need a human label, instead of learning from a randomly labeled set. By querying the examples it is least certain about, it can reach a target accuracy using far fewer labeled examples than random sampling needs.

011 min

The problem active learning solves

A team training a support-ticket classifier has 200,000 unlabeled tickets and a budget for maybe 3,000 labels. The obvious approach is to grab a random 3,000 and pay annotators to label them. That works, but it wastes a lot of the budget on tickets the model would have classified correctly anyway β€” routine "reset my password" tickets that look almost identical to a thousand others already covered. The labels that would actually move the model are the confusing edge cases: tickets that could belong to two categories, or that use unfamiliar phrasing. Random sampling has no way to find those specifically; it just gets whatever it gets.

Labeled data is often the real bottleneck in a machine learning project, not model architecture. Speech recognition is a clear case: annotation at the word level can take ten times longer than the actual audio to produce, and annotating individual sounds, called phonemes, can take up to 400 times as long. Information extraction has the same problem in a different shape β€” locating entities and relations can take a half-hour or more for even simple newswire stories, and specialized domains, such as tagging gene and disease names in biomedical text, need annotators with PhD-level expertise. Active learning exists to spend a fixed labeling budget on the examples that teach the model the most, rather than on whatever happened to be sampled.

021 min

Where active learning comes from

Active learning is not a single algorithm invented on one date by one person β€” it is a subfield that grew out of a simple hypothesis, formalized across many papers through the 1990s and 2000s and gathered into a widely cited synthesis by Burr Settles in his 2009 "Active Learning Literature Survey." The key hypothesis is that if the learning algorithm is allowed to choose the data from which it learns, it will perform better with less training. The same underlying idea appears earlier in statistics under a different name β€” "optimal experimental design" β€” where researchers have long studied how to choose which experiments to run, not just how to analyze results after the fact, in order to learn the most from the fewest trials. Machine learning's active learning is that same design question applied to which data points a model should be trained on.

032 min

How active learning works

The most common setup is pool-based active learning. The model starts with a small labeled set and a large pool of unlabeled data. On each round, it scores every unlabeled example by some measure of how useful its label would be, sends the highest-scoring examples to a human annotator β€” often called an oracle β€” retrains on the newly labeled data, and repeats. The cycle continues until the labeling budget runs out or the model reaches the accuracy it needs.

The simplest and most widely used scoring method is uncertainty sampling: query the example the model is currently least sure how to label. For a classifier, that usually means the example closest to its decision boundary β€” the point where it is essentially guessing between two classes. Other strategies exist, such as query-by-committee, which trains several models and queries the example they disagree on most, but the underlying logic is the same across all of them: spend the labeling budget where the model's current knowledge is weakest, not where it is already confident.

Uncertainty sampling has a known blind spot. The least certain instance lies on the classification boundary, but is not representative of other instances in the distribution, so knowing its label is unlikely to improve accuracy on the data as a whole. A genuine outlier can sit right on the decision boundary without telling the model anything useful about the rest of the data, so a pure uncertainty strategy can waste queries on unrepresentative noise. More advanced strategies build in a measure of how representative a candidate example is of the broader unlabeled pool, specifically to avoid this failure mode.

Query-by-committee takes a different approach to the same problem: train several models β€” the "committee" β€” on the current labeled set, let each one label every unlabeled candidate, and query whichever candidate the committee disagrees on most. An example every committee member already agrees on is unlikely to teach any of them something new; an example that splits the committee is, by definition, one where at least some of the models are wrong, and finding out which one narrows down the space of models that could be correct. Freund et al. (1997) gave this approach a theoretical grounding, not just an intuitive one.

This, like the toy example above, is an exponential improvement over the typical

sample complexity of ordinary supervised learning, where the labeled set is just handed to the model rather than chosen by it. Under the conditions their analysis assumes, the number of labels needed to reach a given error rate can shrink dramatically once the model is allowed to pick its own training examples instead of learning from whatever it is given.

041 min

A worked example

The literature survey that popularized the field includes its own small, checkable demonstration. On a toy two-class dataset, a logistic regression model trained on 30 randomly labeled instances reached only 70% accuracy.

A logistic regression model trained with 30 labeled instances randomly drawn from the problem domain. (70% accuracy)

The same model, trained on the same number of labeled instances β€” but with those 30 chosen by uncertainty sampling instead of at random β€” reached a noticeably higher score.

A logistic regression model trained with 30 actively queried instances using uncertainty sampling (90%).

Twenty percentage points of accuracy came from choosing which 30 points to label, with the label budget held completely fixed. That gap is the entire case for active learning in one small experiment: the same amount of human labeling work produced a meaningfully better model, purely because the queries were chosen rather than random. Real datasets rarely hand over a clean 20-point gain from one toy run, but the direction of the result β€” same label budget, better model, because of which points were chosen β€” is the pattern active learning reliably produces across the far larger and messier datasets it is normally used on.

052 min

Using active learning on a real project

For a machine learning engineer scoping a labeling project, active learning is a direct lever on cost: instead of asking "how many labels can we afford," the question becomes "how much accuracy can we get for a fixed labeling budget," and active learning usually improves that ratio. Popular annotation tools, including Prodigy from the makers of the spaCy NLP library, build active learning suggestions directly into the labeling interface, so an annotator's next task is already chosen by the model's own uncertainty rather than by a random queue.

For a product manager planning a new classifier or extraction feature, active learning changes how to pitch the labeling phase to stakeholders: rather than committing to a large fixed labeling budget upfront, a team can label in small batches, retrain, and check whether accuracy has plateaued before buying the next batch of labels β€” turning an open-ended cost into a measured, stoppable one.

When [a team] wants to [reach a target model accuracy], they want [to spend the labeling budget on the examples the model is least sure about], so they can [get there with far fewer total labels than random sampling would need].

For a founder weighing whether a labeled-data-hungry feature is even feasible on a small budget, active learning is one of the more reliable ways to shrink the number of labels a first working version needs, though it does not remove the need for labels entirely, and it works best when there is a large existing pool of unlabeled data to select from.

Picture the same support-ticket classifier built two ways. In the first version, the team hires annotators to label 3,000 randomly sampled tickets in one batch, trains once, and ships whatever accuracy that produces β€” with no way to know in advance whether 3,000 was enough or twice what was needed. In the second version, the team labels 500 tickets to start, trains a model, has it flag the 500 unlabeled tickets it is least sure about, sends only those for labeling, retrains, and repeats. By the time the second team has spent the same 3,000-label budget, they have typically reached a higher accuracy than the first team, because every batch after the first was chosen to fix the model's current weak points rather than sampled blind. The difference is not the total labels spent β€” it is which 3,000 tickets got labeled.

062 min

When active learning does not help

Active learning's gains are not guaranteed, and the research is explicit that they can go the other way. Guo and Schuurmans (2008) found that off-the-shelf query strategies, when myopically employed in a batch-mode setting are often much worse than random sampling β€” selecting a whole batch of examples at once using a strategy designed for one-at-a-time queries can backΒ­fire, often because the batch ends up full of near-duplicate uncertain examples rather than a diverse, informative set. Results have also been inconsistent across tasks and annotators.

Baldridge and Palmer (2009) found a curious inconsistency in how well active learning helps that seems to be correlated with the proficiency of the annotator β€” a domain expert was better utilized by an active learner than a domain novice, who was better suited to a passive learner, meaning the same query strategy paired with a junior annotator can perform worse than simply handing them a random batch to label.

Active learning also assumes the oracle's labels are trustworthy and that unlabeled data is cheap and plentiful relative to labels β€” assumptions that do not always hold. If the pool of unlabeled data is small, there may not be enough genuinely uninformative examples to skip past, and the savings shrink. And because active learning deliberately selects unusual, borderline examples rather than a representative sample, the resulting labeled set can end up biased toward hard cases, which is exactly what makes it efficient for training but can make it a poor sample to use for estimating a model's overall accuracy afterward β€” that still requires a separately, randomly sampled test set.

072 min

Active learning vs. nearby concepts

Semi-supervised learning also tries to get more value out of a small labeled set, but it does so differently: it uses the unlabeled data directly during training, on the assumption that unlabeled examples still carry useful structure about the data distribution, rather than by asking a human to label more of it. Active learning and semi-supervised learning are often combined β€” label the most informative examples through active learning, and let the model learn from the remaining unlabeled pool through semi-supervised methods β€” but neither technique requires the other.

Reinforcement learning is a different training paradigm entirely: an agent learns by taking actions in an environment and receiving a reward signal, with no human directly labeling individual examples. The confusion mostly comes from vocabulary β€” both fields talk about an agent or model "choosing" what to do next β€” but active learning's choice is about which data point to request a label for for supervised training, while reinforcement learning's choice is about which action to take to maximize a reward over time. A system can use both: an active learning loop to decide what to label, feeding a separate model that is itself trained with reinforcement learning.

Both nearby concepts share active learning's general theme of an algorithm being more selective about what it learns from, which is exactly why the names get mixed up in casual conversation. The test that cuts through the confusion is to ask what is actually being requested at each step: a human-provided label for one specific example (active learning), the structure implicit in unlabeled data with no new human input (semi-supervised learning), or a reward signal for an action just taken (reinforcement learning). Naming that one request correctly usually settles which concept actually applies.

?5 questions

Questions people ask

What is active learning in machine learning?

It is a training approach where the model itself selects which unlabeled examples would be most useful to have labeled next, typically the ones it is currently least certain about, instead of learning from a randomly labeled set.

How much does active learning reduce labeling costs?

It varies by task, but a well-known toy demonstration in Burr Settles' 2009 survey showed a 20-percentage-point accuracy gain (70% to 90%) from the same 30 labels, just by choosing which 30 to label rather than picking them at random.

What is uncertainty sampling?

The most common active learning strategy: on each round, query the label of whichever unlabeled example the current model is least confident about, usually the one closest to its decision boundary between classes.

Does active learning always improve results?

No. Research has found that some query strategies perform much worse than random sampling when applied carelessly in a batch setting, and results can vary depending on the proficiency of the human annotator doing the labeling.

What tools use active learning for data labeling?

Prodigy, an annotation tool built by the makers of the spaCy NLP library, is a well-known example that builds active learning suggestions directly into its labeling workflow.

β–Ά3 videos

Watch

  • Understanding Active Learning

  • Active Learning Overview

    MIT OpenCourseWare

  • What is Active Learning?

    NWIACOMMCOLLEGE

Β§1 source

Sources and further reading

  1. Settles, B. (2009, updated 2010). "Active Learning Literature Survey." Computer Sciences Technical Report 1648, University of Wisconsin–Madison. Full text (PDF).

Keep reading

More from AI

All of AI
All of AI