011 min
Quantization at a glance
- What it is: Storing weights and activations in fewer bits: FP32 down to FP16, INT8, or INT4.
- Why it exists: Model size and inference speed are usually limited by moving data, not by doing arithmetic.
- What it costs: Some accuracy, and naive rounding can break badly on a small set of outlier values.
- When it breaks: Plain 8-bit rounding degrades large transformer models until the outlier problem is handled separately.
- As of: As of 2023, 4-bit quantization let a 65-billion-parameter model be fine-tuned on a single 48GB GPU.
022 min
The problem Quantization solves
Neural networks are usually trained and stored as 32-bit floating point numbers, the format called FP32, because it gives training the most numerical room to work with. Every weight and every activation takes 4 bytes. A model with 7 billion parameters needs about 28 gigabytes just to hold its weights, before counting anything needed to actually run it.
Large language models made this cost impossible to ignore. Dettmers et al. write that "Large language models have been widely adopted but require significant GPU memory for inference." A model with 175 billion parameters stored in FP32 needs roughly 700 gigabytes of memory, far more than any single GPU sold today holds, and often more than one server holds.
The cost is not only how much memory a model needs. It is also how fast that memory can be read. Generating one token with a large model means moving every one of its weights from memory into the processor, then doing so again for the next token. Most of that time goes to moving the weights, not to the arithmetic done on them. This is why running a large model one token at a time is described as memory-bound: cutting the number of bytes that must move cuts the time almost directly.
Quantization addresses both costs at once. Storing each weight in fewer bits shrinks the model on disk and in memory, and it shrinks the number of bytes that must move for every token generated.
032 min
How Quantization works
Take one weight matrix inside a transformer's feed-forward layer, a 4096-by-4096 grid of numbers learned during training. In FP32 that matrix holds about 16.8 million 4-byte numbers, each one a floating point value with roughly seven decimal digits of precision. Quantization replaces every one of those numbers with a small integer, most often an 8-bit integer that can only take one of 256 values.
The conversion works through two numbers computed for the whole matrix, or for a slice of it: a scale factor and a zero-point. To quantize a value x, the scheme computes q = round(x / scale) + zero_point, then clips q into the integer range. For INT8 that range runs from -128 to 127. To use the number again, the same formula runs backward: value β (q - zero_point) Γ scale. The scale is set from the largest and smallest values actually seen in that matrix, a step called calibration, because it decides which real numbers map to which of the 256 available integers.
This is where the interesting step happens. A matrix's actual range of values decides how much precision each integer step represents. If every value in the matrix sits between -2 and 2, calibration can set a tight scale. Each of the 256 integers then stands for a small, precise step. If even one value in that matrix is 200 times larger than the rest, an outlier, the scale must stretch to cover it. Every other value then gets crushed into a handful of integers near zero. One extreme value can destroy the precision of millions of ordinary ones.
There are two ways to decide the scale and zero-point. Post-training quantization (PTQ) calibrates a model that has already finished training. It uses a small batch of sample inputs to see what range of values actually shows up, then converts the weights once. Quantization-aware training instead simulates the rounding during training itself, so the network's own weights adjust around it. Jacob et al. describe this approach: "We also co-design a training procedure to preserve end-to-end model accuracy post quantization." Nagel et al. summarize the trade-off: "PTQ requires no re-training or labelled data and is thus a lightweight push-button approach to quantization. In most cases, PTQ is sufficient for achieving 8-bit quantization with close to floating-point accuracy." QAT costs more, a full or partial retraining pass, but reaches lower bit-widths with less accuracy lost.
Whichever method sets the scale, the benefit at inference time is the same: the 4096-by-4096 matrix now takes a quarter of the memory it did in FP32, and reading it from memory takes a quarter of the time.
042 min
A concrete example: GPTQ
The clearest evidence that quantization works at scale comes from GPTQ, a post-training method published by Frantar et al. in 2022. Before GPTQ, quantizing a model as large as GPT-3's 175 billion parameters to low bit-widths cost either too much accuracy or too much compute to be practical. Frantar et al. report that "GPTQ can quantize GPT models with 175 billion parameters in approximately four GPU hours, reducing the bitwidth down to 3 or 4 bits per weight, with negligible accuracy degradation relative to the uncompressed baseline."
Four GPU-hours is not a typo. A model that took an enormous cluster weeks to train can be quantized to a fraction of its size on a single machine in an afternoon. The result changed what a single GPU could run: the paper states the method was "allowing us for the first time to execute an 175 billion-parameter model inside a single GPU for generative inference."
The gain is not only memory. Frantar et al. also measured how much faster the quantized model ran, finding "end-to-end inference speedups over FP16, of around 3.25x when using high-end GPUs (NVIDIA A100) and 4.5x when using more cost-effective ones (NVIDIA A6000)." The larger speedup on the cheaper GPU fits the memory-bound explanation above: a GPU with less memory bandwidth loses more time waiting for full-precision weights to arrive, so it gains more when quantization cuts the bytes that have to move.
051 min
What this means for your work
For an engineer choosing a model to serve, quantization decides whether a model fits on the hardware available at all, before it decides how fast the model runs. A 13-billion-parameter model that needs about 26 gigabytes in FP16 does not fit on a 24-gigabyte GPU. The same model at INT8 needs about 13 gigabytes, which does fit, with room left for the memory the model uses while actually generating text. The choice of bit-width is often a hardware-fit decision before it is a speed decision.
For a founder or PM setting a product's cost budget, the GPU bill for serving a model scales with how much memory the model occupies and how fast that memory can be read. A model quantized to INT8 or INT4 can often be served on a cheaper GPU tier, or on fewer machines for the same traffic. That is a recurring monthly cost, not a one-time engineering cost, and it compounds every month a product stays live.
For anyone deciding whether to quantize at all, the deciding question is where the current bottleneck sits. A model that is failing on accuracy needs a better model or more training data, not fewer bits. A model whose answers are already good enough, but that is too slow or too expensive to run, is exactly the situation quantization is built for.
061 min
What Quantization costs
The main cost of quantization is accuracy, and it does not fall evenly. Nagel et al. note plainly that "the additional noise it induces can lead to accuracy degradation." For most convolutional and mid-size models, dropping from FP32 to INT8 costs very little. Nagel et al. found that post-training INT8 quantization gets "close to floating-point accuracy" in most cases.
Large transformer models did not follow that pattern. Dettmers et al. found that plain INT8 rounding degraded accuracy sharply once a model passed a few billion parameters. The cause was a small number of feature dimensions inside the model holding unusually large values compared to the rest. The paper names these "highly systematic emergent features in transformer language models that dominate attention and transformer predictive performance." A calibration scale wide enough to cover these outliers crushes every ordinary value into a handful of integers, the same failure the calibration example above describes, just at a scale that breaks the whole model rather than one matrix.
Their fix, called LLM.int8(), "isolates the outlier feature dimensions into a 16-bit matrix multiplication while still more than 99.9% of values are multiplied in 8-bit." That 0.1% accounts for almost all the accuracy risk.
072 min
What Quantization does not solve
Quantization does not fix a slow or inaccurate model. It only removes the memory and bandwidth cost of moving weights between memory and the processor at inference time. It changes nothing about what the model computes or which task it was trained for. A team asking why their quantized 7-billion-parameter model still gives worse answers than a well-trained 3-billion-parameter one is quantizing the wrong problem.
The speedup is also not the same in every setting. It is largest when inference is memory-bound: generating text one token at a time, where the processor spends most of its time waiting for weights to arrive. Training uses higher precision throughout, so quantization does not apply there in the same way. Batched inference, where many requests share the same loaded weights, keeps the processor busier doing arithmetic and spending less time waiting on memory. Quantization then buys less speed for the same drop in precision. Treating "quantization always gives a proportional speedup" as a rule leads teams to quantize a batch-serving pipeline expecting a four-times speedup and finding much less.
Very low bit-widths carry their own limit. GPTQ shows that 3-bit and 4-bit weights can hold accuracy close to the FP16 baseline. Frantar et al. also report that pushing further, to 2-bit or even ternary quantization, is possible but keeps accuracy only "reasonable," not equal to the uncompressed model. Below a certain bit-width, some information cannot be recovered by a better rounding scheme alone; retraining the model around the lower precision, quantization-aware training, is what gets close again.
082 min
Quantization vs. nearby concepts
| Concept | What it does | How it differs from quantization |
|---|---|---|
| AIQuantization-Aware Training | ||
| AIPruning | ||
| AIDistillation |
Three techniques get grouped together as "model compression," and none of them do the same thing.
Quantization-aware training is not a separate idea from quantization. It is one of the two ways of doing quantization: the one that retrains a model around its own rounding instead of rounding it after training finishes. With training data and compute for a retraining pass, it reaches lower bit-widths with less accuracy lost. Without them, post-training quantization is the only option.
Pruning makes a model smaller by removing weights or whole neurons entirely, so the model ends up with fewer numbers. Quantization keeps every number and stores each one in fewer bits. The two are not competitors: a model is often pruned first and quantized after, since they attack different parts of the same memory total.
Distillation does not touch the original model's numbers. It trains a new, smaller model from scratch to copy a larger one's outputs. Quantization compresses a model that already exists; distillation produces a different model that was never the same size to begin with.
092 min
A second case: QLoRA
GPTQ and LLM.int8() both quantize a model to make it cheaper to run. QLoRA, published by Dettmers et al. in 2023, uses quantization to make fine-tuning cheaper instead of making inference cheaper. That is the one variable that differs between the two cases.
The paper describes the core mechanism: "QLoRA backpropagates gradients through a frozen, 4-bit quantized pretrained language model into Low Rank Adapters." The base model's weights are quantized to 4-bit and never updated. Only a small set of added adapter weights, kept in higher precision, are trained. That combination is what makes the memory saving possible: the authors report that QLoRA "reduces memory usage enough to finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance." Without quantizing the frozen base weights, fitting a 65-billion-parameter model's gradients and optimizer state on one GPU would not be possible at all.
Two further details make QLoRA's use of quantization more careful than a single rounding pass. It introduces "4-bit NormalFloat (NF4), a new data type that is information theoretically optimal for normally distributed weights," rather than reusing plain integer quantization, because model weights are not spread evenly across their range. It also applies "double quantization to reduce the average memory footprint by quantizing the quantization constants" themselves, since even the scale and zero-point numbers computed during calibration take up real memory across thousands of matrices.
The result reported in the paper: models fine-tuned this way, named Guanaco, reached "99.3% of the performance level of ChatGPT while only requiring 24 hours of finetuning on a single GPU." The contrast with the first example is the point. GPTQ quantizes a finished model to serve it faster. QLoRA quantizes a model first and trains through the quantization, so the saving shows up during training instead of after it.
102 min
How Quantization changed since
Quantization is much older than deep learning. Compressing numbers into fewer bits is a general idea from digital signal processing, and researchers had applied versions of it to neural networks before 2018. What changed that year was a demonstration that a quantized model's accuracy could be preserved rather than merely tolerated. Jacob et al., working at Google, published a scheme that trained the network with quantization built into the training loop, aimed at running models on phones. They write that their method "improves the tradeoff between accuracy and on-device latency," tested on MobileNets, a model family built for exactly that constraint.
Large language models changed what quantization had to solve. Jacob et al.'s 2018 paper covered mobile-sized vision and detection models with millions of parameters, not the tens or hundreds of billions a language model can carry. When Dettmers et al. quantized large transformers to 8-bit in 2022, they found the outlier-feature problem described earlier, a failure mode that barely shows up at smaller scale and dominates at large scale. GPTQ, the same year, pushed post-training quantization down to 3 and 4 bits at 175-billion-parameter scale, something the 2018 methods were never tested against. QLoRA, in 2023, took quantization somewhere none of the earlier work went: inside the training loop of a large model, rather than only at the end of it.
The throughline is scale: every later paper solved a problem the 2018 techniques only ran into once models grew far past what existed when they were published.
112 min
Frequently asked questions about Quantization
Is quantization the same as pruning?
No. Quantization stores every existing weight in fewer bits. Pruning removes weights or neurons entirely, leaving fewer numbers behind. A model can go through both: prune first to cut the parameter count, then quantize what remains to shrink each remaining number.
When should I use post-training quantization instead of quantization-aware training?
Use post-training quantization when there is no budget for retraining and 8-bit accuracy is enough. Use quantization-aware training when you need 4-bit or lower and have labeled data and compute to retrain the model around the rounding.
What are the risks of quantizing a model?
The main risk is accuracy loss, and it is uneven rather than gradual. Large transformer models can degrade sharply at 8-bit unless outlier feature values get special handling, the problem LLM.int8() was built to solve, while smaller models often lose almost nothing.
Does quantization require a fine-tune?
Not always. Post-training quantization needs no retraining, just a small calibration set of sample inputs. Quantization-aware training and methods like QLoRA do involve training, either to adapt the model to rounding or to fine-tune it through frozen quantized weights.
How much memory does quantization save?
Moving from 32-bit to 8-bit cuts memory to a quarter. Moving to 4-bit cuts it to an eighth. GPTQ reports quantizing a 175-billion-parameter model to 3 or 4 bits with inference speedups of roughly 3.25 to 4.5 times over FP16, depending on the GPU.
Can a quantized model run on a CPU?
Yes, and this is one of the main reasons INT8 quantization exists. Integer arithmetic runs efficiently on ordinary CPU hardware, which is why Jacob et al.'s 2018 paper aimed quantization specifically at integer-only, GPU-free inference for mobile devices.
What is 4-bit quantization?
4-bit quantization stores each weight as one of 16 possible values instead of the 256 values 8-bit allows. QLoRA's NF4 format and GPTQ both quantize to 4 bits, cutting memory to an eighth of FP32 with only a small, measured accuracy cost.
?7 questions
Questions people ask
Is quantization the same as pruning?
When should I use post-training quantization instead of quantization-aware training?
What are the risks of quantizing a model?
Does quantization require a fine-tune?
How much memory does quantization save?
Can a quantized model run on a CPU?
What is 4-bit quantization?
Β§5 sources
Sources
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. arXiv:1712.05877.
Dettmers, T., Lewis, M., Belkada, Y., & Zettlemoyer, L. (2022). LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. arXiv:2208.07339.
Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2022). GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. arXiv:2210.17323.
Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. arXiv:2305.14314.
Show all 5 sourcesShow fewer sources
Nagel, M., Fournarakis, M., Amjad, R. A., Bondarenko, Y., van Baalen, M., & Blankevoort, T. (2021). A White Paper on Neural Network Quantization. arXiv:2106.08295.





