Distillation

A technique where a small model is trained to copy the behavior of a large one, keeping most of the accuracy while being cheaper and faster to run.

9 min read

Β· Also in

By Ravi SuranaUpdated 3 sources

Quick answer

~20 sec

Distillation is a way to shrink an AI model. A large, accurate "teacher" model is used to train a small "student" model to copy its outputs. The student ends up far cheaper and faster to run while keeping most of the teacher's accuracy. It is how large models get compressed to run on phones or under tight budgets.

011 min

The problem distillation solves

The most accurate models are often the most expensive to run. The naive way to get top accuracy is to train a very large neural network, or an ensemble of several models whose predictions are averaged together. Both approaches work, and both are painful to deploy. Making a prediction from an ensemble of large models is slow and computationally expensive, and pushing that cost onto every user request β€” for a feature used millions of times a day β€” is often impossible within a real latency and hardware budget.

The obvious alternatives each give something up. You can just train a small model directly, but a small model trained from scratch on the raw labels usually lands well short of the large one's accuracy. You can serve the large model and eat the cost, but that caps how many users you can serve and rules out running on a phone or a browser entirely. Neither is satisfying when you want the accuracy of the big model and the cost of the small one.

Distillation was designed for exactly this gap: a way to compress the knowledge in a large, expensive model into a small model that is much easier to deploy, without simply retraining the small model on the original labels and accepting the accuracy hit.

022 min

How distillation works

Follow one training example: a photo the large teacher model classifies. Trained normally, a small student would only see the hard label β€” "this is a dog" β€” and try to reproduce it. Distillation gives the student more than that. It trains the student to match the teacher's full output: not just "dog," but the teacher's probability across every class β€” say 90% dog, 8% wolf, 2% cat.

That richer target is the whole trick. The teacher's soft probabilities carry information the hard label throws away: they say a dog looks somewhat like a wolf and almost nothing like a cat. This is the "dark knowledge" in the teacher β€” its learned sense of how classes relate β€” and the student learns far more from copying the whole distribution than from the one-word answer alone. In practice the probabilities are softened with a temperature setting that exaggerates the small differences between the runner-up classes, so the student can see them clearly.

The student is trained on a combination of two signals: matching the teacher's softened outputs (the distillation signal) and getting the real labels right (the ordinary signal). The result is a small model that has absorbed the teacher's relational knowledge, not just its final answers. Distillation is closely related to fine-tuning in that both adapt a model with further training, but the goal is different: fine-tuning specializes a model for a task, while distillation transfers one model's behavior into a smaller one.

031 min

A concrete example

The clearest checkable case is DistilBERT, a distilled version of the BERT transformer language model released by Hugging Face researchers in 2019. BERT was a large, accurate model that was expensive to run, which made it awkward for on-device or budget-constrained use. The team distilled it during pre-training and reported a specific, measurable result: DistilBERT reduces the size of a BERT model by 40%, while retaining 97% of its language understanding capabilities and being 60% faster.

Those three numbers are what make distillation concrete rather than hopeful. A 40% smaller model that keeps 97% of the capability and runs 60% faster is a trade most teams will take gladly, because the 3% accuracy give-up buys a model that fits where the original could not. The result also shows what distillation is not: it is not lossless. The student is measurably a little worse than the teacher, and the engineering question is always whether the size and speed win is worth that specific, quantified accuracy cost for the task at hand. For a search feature that runs on every keystroke, or a model that has to run inside a browser tab, a 60% speedup for a 3% accuracy give-up is not a close call β€” it is the difference between a feature that ships and one that does not.

041 min

A second case: the original distillation results

DistilBERT shows distillation compressing one big model into a smaller one. The 2015 paper that named the technique showed a different shape of the same idea, and the contrast is instructive. Hinton, Vinyals, and Dean applied distillation to compress an ensemble of models β€” many models whose predictions are averaged β€” into a single model. They reported surprising results on MNIST, the standard handwritten-digit benchmark, and, more consequentially, showed they could significantly improve the acoustic model of a heavily used commercial speech system by distilling the knowledge in an ensemble of models into a single model.

The variable that differs from the DistilBERT case is what is being compressed. DistilBERT compresses one large model into a smaller one; the original speech result compresses many models into one. The contrast teaches that "teacher" does not have to mean a single big network β€” it can be an ensemble whose combined judgment is what the student absorbs. That is why distillation is described as transferring knowledge rather than just shrinking a model: the student can inherit the pooled behavior of several teachers, which no single one of them could have supplied alone, while costing no more to run than one small model.

051 min

What distillation means for your work

For a product manager, distillation is what makes an expensive model's accuracy affordable at scale. If a feature works with a large model in a demo but the per-request cost or latency makes it impossible to ship to every user, distillation is often the path from prototype to production. The concrete decision it changes: rather than cutting the feature or accepting a weaker off-the-shelf small model, the plan becomes "ship the big model to a subset, use it to distill a small one, then serve the small one" β€” which turns an unaffordable feature into an affordable one.

For an engineer, the practical appeal is that distillation needs the teacher's outputs, not its internals or its original training data. You can often distill from a model you can query even when you cannot retrain it, using its predictions on your own unlabeled data as the training signal. This is why distillation shows up constantly in deployment work β€” compressing a large language model into something that fits a latency budget, or specializing a general model into a fast one for a single task.

For a founder, distillation is part of why a small team can serve a capable model cheaply: you do not always need to own the biggest model, only a small one that has learned to imitate a bigger one well enough for your specific use.

061 min

What distillation costs

Distillation is not free, and its costs land in a specific order. The upfront cost is training: you still need the large teacher model, either to train or to query, and you pay to run it across a large body of examples to generate the soft targets the student learns from. For a big teacher and a lot of data, generating those targets is itself a meaningful compute bill, paid once at training time.

The payoff is at inference, which is where cost usually matters most. The student is smaller and faster for every request forever after, so a one-time training expense buys an ongoing per-request saving β€” the reason distillation is worth it for high-traffic features and rarely worth it for something queried a handful of times.

The quieter cost is the accuracy give-up and its maintenance. The student is always at least slightly worse than the teacher, and if the teacher improves or the data drifts, the student has to be re-distilled to keep up. A distilled model is a snapshot of its teacher at one moment, not a living copy that improves alongside it.

071 min

What distillation does not solve

Distillation cannot give a student capabilities the teacher never had. It transfers what the teacher knows; it does not add knowledge. If the teacher is wrong about something, the student faithfully learns to be wrong in the same way, so distillation is a poor fix for a model whose problem is accuracy rather than cost.

It is also the wrong tool when the student is simply too small to hold the teacher's knowledge. There is a floor: past a point, shrinking the student further stops preserving the behavior and the accuracy collapses, because a model with too few parameters cannot represent what the teacher learned however good the training signal is. Distillation narrows the gap between a small model and a large one; it does not erase the fact that capacity has limits.

The overclaim to correct is that a distilled model "is" the big model in miniature. It is not. It is a smaller model trained to imitate the big one's outputs on the data it was distilled with, and it can diverge on inputs unlike that data. Treating the student as a perfect stand-in for the teacher, rather than a cheaper approximation with measurable gaps, is where distillation projects go wrong.

081 min

Distillation vs. nearby concepts

Distillation is one of several model-compression techniques and is easy to confuse with the others. Quantization shrinks a model by storing its numbers at lower precision β€” using 8-bit integers instead of 32-bit floats, for example. It changes how the same model is stored and computed; it does not train a new, smaller model. The deciding difference: quantization keeps the model's structure and reduces the precision of its weights, while distillation trains a genuinely smaller model with fewer parameters.

Pruning is a third technique: it removes weights or whole neurons judged unimportant from an existing model, leaving a sparser version of the same network. Distillation, by contrast, starts a fresh small model and trains it to imitate the large one. The three are often combined β€” a model can be distilled, then pruned, then quantized β€” because they attack the cost from different angles.

Distillation also differs from ordinary fine-tuning. Fine-tuning continues training one model to specialize it for a task; distillation trains a second, smaller model to copy a first one. The signal is the key difference: fine-tuning learns from task labels, while distillation learns primarily from another model's outputs.

091 min

How distillation changed since

The technique was formalized in a 2015 paper, "Distilling the Knowledge in a Neural Network," by Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, which built on earlier model-compression work by Caruana and collaborators. Their framing β€” a small model trained on the softened outputs of a larger one β€” is still the standard recipe, and they showed it working on digit recognition and on a heavily used commercial speech system.

What changed since is scale and ubiquity. Distillation moved from a compression trick to a routine step in shipping models, especially for language. DistilBERT in 2019 made the case concretely for transformers, and as large language models grew, distilling a big model into a smaller, cheaper one became a standard part of the deployment pipeline rather than a research option. The most recent shift, as of 2026, is that distillation is increasingly used not just to compress a model but to transfer capability from a frontier model into a smaller open one β€” a practice powerful enough that model providers now write terms of service specifically about whether their outputs may be used to train competing models.

?8 questions

Questions people ask

What is knowledge distillation in simple terms?

It is training a small "student" model to copy the outputs of a large "teacher" model. The student ends up much cheaper and faster to run while keeping most of the teacher's accuracy, which is how big models get compressed for phones or tight budgets.

How is distillation different from just training a small model?

A small model trained on raw labels only sees the final answer. Distillation trains it on the teacher's full probability distribution, which carries extra information about how classes relate. The student learns far more from copying the whole distribution than from the one-word label.

Who invented knowledge distillation?

It was formalized in a 2015 paper by Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, building on earlier model-compression work by Caruana and colleagues. Their teacher-student recipe using softened outputs is still the standard approach.

What is the difference between distillation and quantization?

Quantization stores the same model's numbers at lower precision, keeping its structure. Distillation trains a genuinely smaller model with fewer parameters to imitate the large one. They are often combined, since they cut cost in different ways.

Does distillation lose accuracy?

Yes, a little. The student is always at least slightly worse than the teacher. DistilBERT, for example, retained 97% of BERT's language understanding while being 40% smaller β€” the small accuracy give-up buys a big size and speed win.

How much smaller and faster can a distilled model be?

It depends on the task, but DistilBERT is a concrete benchmark: 40% smaller and 60% faster than BERT while keeping 97% of its capability. Larger reductions are possible but the accuracy give-up grows as the student shrinks.

Do you need the teacher's training data to distill it?

Not necessarily. You need the teacher's outputs, which you can generate by running it on your own unlabeled data. This lets you distill from a model you can query even when you cannot access its original training set or internals.

Can a distilled model be better than its teacher?

Generally no. Distillation transfers what the teacher knows and cannot add capabilities it never had. If the teacher is wrong, the student learns the same mistake. It narrows the gap between small and large models but does not erase capacity limits.

Β§3 sources

Sources and further reading

  1. Hinton, G., Vinyals, O., & Dean, J. (2015). "Distilling the Knowledge in a Neural Network." arXiv:1503.02531. The paper that formalized knowledge distillation.

  2. Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). "DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter." arXiv:1910.01108. The 40%-smaller, 97%-retained, 60%-faster result.

  3. Buciluă, C., Caruana, R., & Niculescu-Mizil, A. (2006). "Model compression." The earlier work on compressing an ensemble into a single model that the Hinton paper builds on and credits.

Keep reading

More from AI

All of AI
All of AI