Pruning

Pruning is removing weights, neurons, or layers from a trained neural network to shrink it while keeping accuracy close to the original.

10 min read

Β· Also in

By Ravi SuranaUpdated 6 sources

Quick answer

~20 sec

Pruning is removing weights or entire neurons from a trained neural network to make it smaller and cheaper to run, while trying to keep its accuracy close to the original. It works because most trained networks carry far more parameters than the task actually needs. A pruned network is typically fine-tuned afterward to recover any accuracy the removal cost.

011 min

Pruning at a glance

  • What it is: Removing weights, neurons, or whole layers from a trained network to shrink it, on the finding that most networks are trained with far more capacity than the task strictly needs.
  • Why it exists: A full-size network can be too large, slow, or power-hungry to run on a phone, a browser, or a real-time budget, even when its accuracy is exactly what you want.
  • What it costs: Pruning too aggressively without retraining destroys accuracy; recovering it usually costs a full fine-tuning pass, which is compute and engineer time, not a free discount.
  • As of 2023, the technique reaches large language models too β€” methods like Wanda prune LLMs without retraining at all, something older pruning approaches could not do.

021 min

The problem Pruning solves

Before pruning, the standard fix for a network that was too slow or too large to deploy was to train a smaller network from scratch. The 1990 paper that introduced pruning as a formal technique, Yann LeCun, John Denker, and Sara Solla's "Optimal Brain Damage," names the same problem this library sees in the AI category generally: bigger networks were becoming standard because they trained more reliably, but nobody had a principled way to make a trained network smaller afterward without just guessing at a new architecture and starting over.

Training a smaller network from scratch has a real cost beyond wasted compute: it is not guaranteed to reach the same accuracy, because a smaller architecture can fail to find a good solution during training even when a similarly-sized solution exists inside a larger, already-trained network. Pruning sidesteps the guess by starting from a network that already works and removing only the parts that turn out not to matter.

032 min

How Pruning works

LeCun, Denker, and Solla state the basic mechanism plainly: "By removing unimportant weights from a network, several improvements can be expected: better generalization, fewer training examples required, and improved speed of learning and/or classification." Their method, Optimal Brain Damage, used second-derivative information from the loss function to estimate which weights, if deleted, would raise the network's error the least β€” a more principled version of the simpler approach that dominates practice today, magnitude pruning, which just removes the weights with the smallest absolute value on the reasoning that a weight near zero is barely influencing the network's output anyway.

Follow one network the whole way through to see how this plays out. Take an image classifier trained to recognize handwritten digits, with a million weights connecting its layers. After training, most of those weights turn out to be small: many neurons respond to noise as much as signal, and many connections carry almost no information once the network has settled into a good solution. A magnitude-pruning pass sorts all million weights by absolute value and removes, say, the smallest 80%, snapping them to exactly zero. The network's accuracy usually drops immediately after this cut, because even small weights were contributing something. A short fine-tuning pass β€” continuing to train the now-sparse network on the same data β€” lets the surviving 20% of weights adjust to compensate, and accuracy typically recovers close to its original level.

Two structural choices flex this basic loop. Unstructured pruning removes individual weights wherever they fall, which achieves the highest compression ratio but produces an irregular sparsity pattern that ordinary hardware cannot skip over efficiently. Structured pruning removes whole neurons, filters, or channels instead, which compresses less aggressively per parameter removed but produces a smaller dense network that runs faster on standard hardware without any special support. And pruning is rarely a single cut: iterative pruning repeats prune-then-fine-tune in small steps β€” removing 10-20% at a time rather than 80% at once β€” which consistently reaches a higher final sparsity at the same accuracy than one large cut does.

041 min

A concrete example

The clearest real numbers come from Song Han, Huizi Mao, and William Dally's 2015 "Deep Compression" paper, which combined pruning with two other compression steps and measured the result directly. In their own words: "Pruning, reduces the number of connections by 9Γ— to 13Γ—; Quantization then reduces the number of bits that represent each connection from 32 to 5." Chained together with a final entropy-coding step, "our method reduced the storage required by AlexNet by 35Γ—, from 240MB to 6.9MB, without loss of accuracy," and reduced VGG-16 from 552MB to 11.3MB, a 49x reduction, again without an accuracy loss the paper reports.

ModelOriginal sizeAfter Deep CompressionReduction
AlexNet240 MB6.9 MB35Γ—
VGG-16552 MB11.3 MB49Γ—

Pruning alone accounted for the 9-13x of that combined result; quantization and coding did the rest. That split matters for a practical reason: pruning and quantization attack different kinds of redundancy β€” pruning removes weights that barely matter, quantization stores the weights that remain more cheaply β€” which is why production compression pipelines usually apply both rather than picking one.

052 min

What Pruning means for your work

For a founder or PM deciding whether a model can ship on-device, pruning is the difference between "runs only on a server we pay for per request" and "runs on the phone the user already owns," changing the unit economics of a feature rather than just its speed. A model that must live inside a mobile app's size budget, or run inference without a network round-trip, often has to be pruned (or quantized, or both) before it fits at all β€” not as an optimization pass at the end, but as a requirement decided before the feature is scoped.

For an engineer building an inference pipeline, the practical warning is that unstructured pruning's compression numbers do not automatically translate into faster inference. A network with 90% of its weights zeroed out still occupies the same dense matrix shape on most standard GPUs unless the runtime specifically supports sparse computation β€” the FLOPs saved on paper are not always FLOPs saved in wall-clock time. Structured pruning, which removes whole channels rather than scattered individual weights, is the safer default when the deployment target has no special sparse-matrix support.

For an AI or eval team working with large language models specifically, pruning changed shape again in 2023. Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter's Wanda method prunes pretrained LLMs by weight magnitude multiplied by the corresponding input activation, and the paper states plainly that "Wanda requires no retraining or weight update, and the pruned LLM can be used as is." That is a meaningful departure from the Deep Compression-era assumption that pruning always needs a fine-tuning pass afterward β€” for a model with billions of parameters, skipping that pass is often the difference between a technique that is actually affordable to run and one that is not.

061 min

What Pruning costs

The two costs that matter most are compute and accuracy, and they trade against each other directly. Iterative pruning with fine-tuning between each step is more expensive to run than a single cut, in proportion to how many rounds it takes β€” but a single aggressive cut without fine-tuning between rounds reliably loses more accuracy than the same total sparsity reached gradually. Deep Compression's own pipeline needed a retraining pass after pruning specifically to recover the accuracy the cut cost.

The second cost is easy to miss: a compression ratio on paper is not a latency win in production unless the hardware and inference runtime actually exploit the resulting sparsity pattern. Unstructured pruning can report a 10x parameter reduction that produces close to 0x real speedup on hardware without sparse-matrix support β€” the saved parameters are real, the saved milliseconds are not, unless the whole stack downstream is built to use them.

071 min

What Pruning does not solve

Pruning does not fix a model that is fundamentally the wrong architecture for the task, or fundamentally under-trained β€” it only removes redundancy from a network that has already learned something worth keeping. Applying it to a barely-trained network mostly removes weights that had not yet had a chance to become useful, and fine-tuning afterward is really just continuing an interrupted training run under a misleading name.

The 2020 benchmark study "What is the State of Neural Network Pruning?" by Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag makes a second limitation explicit at the level of the whole field, not one method: after aggregating results from 81 papers, the authors report that "the community suffers from a lack of standardized benchmarks and metrics," and that "this deficiency is substantial enough that it is hard to compare pruning techniques to one another or determine how much progress the field has made over the past three decades." A claimed state-of-the-art pruning result should be read with that caveat attached β€” the comparison it is claiming to beat may not have been measured the same way.

081 min

Pruning vs. nearby concepts

The nearest neighbours are quantization and distillation, and all three get grouped together as "model compression" loosely enough that the differences blur. Pruning removes weights entirely, leaving the remaining ones at full precision. Quantization keeps every weight but stores each one with fewer bits, trading numeric precision for size. Distillation trains an entirely new, smaller network from scratch to copy a larger one's behaviour, rather than editing the large network at all. Deep Compression's own pipeline is the clearest illustration that these are not competitors: it applies pruning first, then quantization, then a final coding step, because each targets a different kind of redundancy and the gains multiply rather than overlap.

A related but distinct research question is whether a pruned subnetwork could have been trained on its own from the start. Jonathan Frankle and Michael Carbin's 2019 "lottery ticket hypothesis" paper found that standard pruning "naturally uncovers subnetworks whose initializations made them capable of training effectively" β€” small "winning ticket" subnetworks that train to comparable accuracy in isolation. That is a claim about why pruning works, not a competing compression technique, and it is contested β€” see below.

091 min

Where the evidence is contested

The lottery ticket hypothesis's original claim, tested mainly on small networks on MNIST and CIFAR-10, did not hold up unchanged at larger scale. A 2020 follow-up by Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin, published at ICML, found that the winning subnetworks the method identifies "only reach full accuracy when they are stable to SGD noise, which either occurs at initialization for small-scale settings (MNIST) or early in training for large-scale settings (ResNet-50 and Inception-v3 on ImageNet)." In plain terms: for the larger, more realistic networks the field actually cares about, the winning ticket has to be identified a little way into training, not at random initialization as the original, simpler version of the hypothesis proposed.

This is a genuine, unresolved complication rather than a debunking β€” the same authors, in the same line of research, are the ones who found and reported the limitation. Where it leaves a practitioner: the lottery ticket framing is a useful way to think about why iterative magnitude pruning tends to work, but "prune from a random initialization and get a trainable small network for free" is not the reliable recipe the earliest framing suggested, and reproducing it at the scale of a modern production model needs the early-training checkpoint the 2020 paper identifies, not the initial random weights.

102 min

Frequently asked questions about Pruning

What is pruning in machine learning?

It's removing weights, neurons, or layers from a trained neural network to make it smaller and faster, based on the finding that most trained networks carry more parameters than the task strictly needs.

Does pruning hurt model accuracy?

A pruning cut usually drops accuracy immediately, but fine-tuning the remaining weights afterward typically recovers most or all of the loss, up to a point β€” pruning too aggressively without enough fine-tuning degrades accuracy permanently.

What is the difference between pruning and quantization?

Pruning removes weights entirely, leaving the rest at full precision. Quantization keeps every weight but stores each with fewer bits. They target different redundancy and are usually combined, not chosen between.

Does pruning require retraining?

Traditional pruning does, to recover accuracy lost in the cut. Newer methods built for large language models, like Wanda (2023), are designed to skip retraining entirely, because retraining a billion-parameter model is often unaffordable.

How much can pruning shrink a model?

The 2015 Deep Compression paper reduced AlexNet's storage 35x and VGG-16's 49x when pruning was combined with quantization and coding; pruning alone accounted for a 9x to 13x reduction in connections in that result.

Does pruning always make inference faster?

Not automatically. Unstructured pruning can report a large parameter reduction with little real speedup unless the hardware and runtime specifically support sparse computation β€” structured pruning, which removes whole channels, is the safer default for ordinary hardware.

?6 questions

Questions people ask

What is pruning in machine learning?

It's removing weights, neurons, or layers from a trained neural network to make it smaller and faster, based on the finding that most trained networks carry more parameters than the task strictly needs.

Does pruning hurt model accuracy?

A pruning cut usually drops accuracy immediately, but fine-tuning the remaining weights afterward typically recovers most or all of the loss, up to a point β€” pruning too aggressively without enough fine-tuning degrades accuracy permanently.

What is the difference between pruning and quantization?

Pruning removes weights entirely, leaving the rest at full precision. Quantization keeps every weight but stores each with fewer bits. They target different redundancy and are usually combined, not chosen between.

Does pruning require retraining?

Traditional pruning does, to recover accuracy lost in the cut. Newer methods built for large language models, like Wanda (2023), are designed to skip retraining entirely, because retraining a billion-parameter model is often unaffordable.

How much can pruning shrink a model?

The 2015 Deep Compression paper reduced AlexNet's storage 35x and VGG-16's 49x when pruning was combined with quantization and coding; pruning alone accounted for a 9x to 13x reduction in connections in that result.

Does pruning always make inference faster?

Not automatically. Unstructured pruning can report a large parameter reduction with little real speedup unless the hardware and runtime specifically support sparse computation β€” structured pruning, which removes whole channels, is the safer default for ordinary hardware.

Β§6 sources

Sources

  1. LeCun, Y., Denker, J. S., & Solla, S. A. (1990). Optimal Brain Damage. Advances in Neural Information Processing Systems 2.

  2. Han, S., Mao, H., & Dally, W. J. (2016). Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. ICLR 2016.

  3. Frankle, J., & Carbin, M. (2019). The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks. ICLR 2019.

  4. Frankle, J., Dziugaite, G. K., Roy, D. M., & Carbin, M. (2020). Linear Mode Connectivity and the Lottery Ticket Hypothesis. ICML 2020.

Show all 6 sources
  1. Blalock, D., Gonzalez Ortiz, J. J., Frankle, J., & Guttag, J. (2020). What is the State of Neural Network Pruning? MLSys 2020.

  2. Sun, M., Liu, Z., Bair, A., & Kolter, J. Z. (2023). A Simple and Effective Pruning Approach for Large Language Models (Wanda). ICLR 2024.

Keep reading

More from AI

All of AI
All of AI