Backpropagation

The algorithm that trains a neural network by computing how much each weight contributed to a wrong output, then adjusting every weight to reduce that error.

10 min read

· Also in

By Ravi SuranaUpdated

Quick answer

~20 sec

Backpropagation is the algorithm that trains a neural network by working backward from its error: it calculates how much each individual weight contributed to a wrong output, then nudges every weight slightly in the direction that would have reduced that error. It is how a network "learns" from thousands of examples, one small adjustment at a time.

011 min

Key takeaways

  • What it is: the backward pass that computes how to adjust every weight in a neural network
  • Why it exists: networks with many layers have no direct way to know which weight caused an error
  • What it costs: a full backward pass roughly doubles the compute of the forward pass alone
  • As of: still the algorithm behind essentially every deep neural network trained in September 2026

021 min

The problem Backpropagation solves

Before backpropagation, the only network anyone could reliably train was the perceptron, a single layer of weights connecting inputs directly to an output, using a simple rule that adjusted each weight based only on that weight's own direct contribution to the final error. That worked because a single-layer network has no intermediate steps to account for: every weight touches the output directly.

The moment a network gains a hidden layer, a layer of units sitting between the input and the output, the direct rule breaks down. A hidden unit's weight does not touch the final output directly; it touches another unit, which touches another, which eventually produces the output several steps later. Nobody had an efficient way to work out how much a weight several layers back from the output should be blamed for a wrong answer, so multi-layer networks stayed impractical to train even though the mathematics of a single layer was well understood. Minsky and Papert's 1969 book Perceptrons documented exactly this limit, showing that a single-layer perceptron could not even learn a function as simple as exclusive-or, and the absence of a practical way to train deeper networks that could learn it contributed to a long stretch, later called an AI winter, when funding for neural network research largely dried up.

032 min

How Backpropagation works

Take a tiny network built to guess whether a support ticket is urgent, from two numeric signals: how many exclamation marks it contains, and how many hours since the account last logged in. The network has one hidden layer of two units between the input and a single output unit that reports a score from 0 to 1. Every connection between units carries a weight, a number that scales how strongly one unit's output affects the next unit's input.

Training starts with a forward pass: the two input numbers are multiplied by their weights, summed, passed through an activation function at each hidden unit (a function that decides how strongly a unit fires given its summed input), and the process repeats going into the output unit, which produces a score. Say the network outputs 0.2 for a ticket a human labelled as genuinely urgent, which should have scored close to 1. The gap between 0.2 and 1 is the error, measured by a loss function.

Backpropagation is the step where the network works out how to fix that error, moving backward through exactly the same connections the forward pass used. Using the chain rule, a rule from calculus for finding how a small change in one variable ripples through a chain of dependent variables, the algorithm computes the gradient: for every single weight in the network, how much that one weight's small increase or decrease would have changed the final error. The output layer's weights are easiest, since they touch the error directly. The hidden layer's weights are harder, since their effect on the error only exists because of how they feed into the output layer, and the chain rule is exactly the tool for multiplying those effects together across layers. Once every weight has a gradient, each weight is nudged a small step in the direction that would reduce the error, a step controlled by a setting called the learning rate. Repeat the forward pass and the backward pass over thousands of labelled tickets, and the same two hidden units gradually settle on weights that separate urgent tickets from ordinary ones.

041 min

A concrete example

The clearest demonstration of what backpropagation actually buys a multi-layer network is the exact limit Minsky and Papert had described: a single-layer perceptron cannot learn exclusive-or, a function that outputs true when exactly one of two binary inputs is true, because no single straight-line boundary can separate its true cases from its false ones. A network with one hidden layer, trained with backpropagation, learns exclusive-or reliably, because the hidden layer can bend the boundary into a shape a single layer never could.

An experimental analysis of the known method by Rumelhart et al. then demonstrated that backpropagation can yield useful internal representations in hidden layers of NNs. The hidden units did not just pass information through unchanged; they organised themselves during training to represent features of the problem nobody had hand-designed, which is exactly the capability multi-layer networks had been missing. A Transformer-based model trained today runs the identical backward-pass procedure at a scale Rumelhart, Hinton and Williams never tested: a model with tens of billions of parameters, built almost entirely from the self-attention layers a transformer uses to weigh which earlier tokens matter to each new one, still computes one gradient per weight, using the same chain rule, on every training step, just with the bookkeeping handled automatically rather than derived by hand for each new network shape.

052 min

What this means for your work

An engineer choosing a deep learning framework almost never writes backpropagation by hand. In PyTorch, the backward pass kicks off when .backward() is called on the DAG root, and autograd computes the gradients from each grad_fn, accumulating them in each tensor's grad attribute, using the chain rule, all the way back to the leaf tensors. Knowing what that call is actually doing changes how an engineer debugs a model that refuses to learn: a loss that never decreases often means a gradient is not flowing backward through some part of the network at all, not that the data or the architecture choice was wrong.

For a founder or product manager deciding whether to fine-tune a model versus prompt it, the relevant fact is that fine-tuning runs the same backward pass the original training run used, just on a smaller dataset and usually through only a subset of the model's weights, which is why fine-tuning is far cheaper than training from scratch but still needs the same gradient machinery and the same GPU memory to store those gradients. An AI evaluator reading a training log for the first time benefits from the same mental model: a vanishing gradient, a gradient that becomes too small to matter by the time it reaches an early layer, is not a mysterious failure but a direct, measurable consequence of how many multiplications the chain rule performs to reach that layer.

061 min

What Backpropagation costs

A full backward pass costs roughly the same amount of compute as the forward pass that produced the prediction, so training a model costs on the order of two to three times what running it afterward costs per example, once the extra bookkeeping is included. The real expense is memory, not arithmetic: computing every weight's gradient requires holding onto the intermediate values from the forward pass until the backward pass needs them, which is why training a large model needs far more GPU memory than simply running the same model to get an answer, and why techniques like gradient checkpointing trade extra recomputation for lower memory use. This cost scales with every parameter in the network, which is the direct reason a model with more weights costs more to train, independent of how much data it sees.

071 min

What Backpropagation does not solve

Backpropagation computes gradients correctly whenever a network is built entirely out of differentiable operations, functions smooth enough to have a well-defined slope everywhere, but it has nothing to say about a network that includes a genuinely non-differentiable step, such as a hard decision to select one of several discrete actions. Reinforcement learning systems that must make a discrete choice work around this with separate machinery, such as policy gradients, rather than backpropagating through the choice itself.

A related, well-known problem is vanishing and exploding gradients: in a very deep network, the same chain-rule multiplication that lets backpropagation reach every layer can shrink gradients toward zero or blow them up toward infinity by the time they reach the earliest layers, which is why plain deep networks were hard to train reliably until architectural fixes such as residual connections and careful weight initialisation became standard. Backpropagation is also not a general-purpose optimiser; it only computes the gradient. A separate algorithm, such as stochastic gradient descent or Adam, decides how big a step to actually take using that gradient, and a poor choice there can stall training even when every gradient is computed correctly.

081 min

Backpropagation vs. nearby concepts

ConceptWhat it namesHow it differs
BackpropagationThe method for computing the gradient of every weight in a network efficientlyAnswers how much to blame each weight
Gradient DescentThe rule for updating a weight once its gradient is knownAnswers how far to move each weight, using the gradient backpropagation already computed
Activation FunctionThe function inside each unit that decides how strongly it firesWhat backpropagation differentiates through at every layer, not the algorithm itself
Autograd (in a framework like PyTorch)Software that implements backpropagation automatically for any network built from its operationsAn implementation of backpropagation, not a different algorithm

A common confusion collapses backpropagation and gradient descent into one word. Backpropagation only computes the gradient; gradient descent, or a variant such as Adam, is the separate step that decides how big an update to actually make with that gradient.

091 min

How Backpropagation changed since

Jürgen Schmidhuber's widely cited 2015 survey states: "I review deep supervised learning (also recapitulating the history of backpropagation), unsupervised learning, reinforcement learning & evolutionary computation, and indirect search for short programs encoding deep and large networks." That survey traces backpropagation's origin further back than most accounts credit. Explicit, efficient error backpropagation (BP) in arbitrary, discrete, possibly sparsely connected, NN-like networks was first described in a 1970 master's thesis (Linnainmaa, 1970, 1976). That method, also known as the reverse mode of automatic differentiation, predates its use in neural networks by more than fifteen years. Paul Werbos discussed a related idea in his 1974 PhD thesis and published the first neural-network-specific application in 1982, also without reaching the researchers who later popularised the technique.

David Rumelhart, Geoffrey Hinton and Ronald Williams's 1986 Nature paper is the one most often credited with inventing backpropagation. This experimental analysis of backpropagation did not cite the origin of the method, also known as the reverse mode of automatic differentiation. It nonetheless demonstrated something the earlier work had not: that the algorithm, applied to real networks, produced hidden units that organised themselves into useful internal representations, which is what actually convinced the field the technique was worth adopting widely. The mathematics the 1986 paper popularised had already been discovered, and largely ignored, twice before.

?6 questions

Questions people ask

Is backpropagation the same as gradient descent?

No. Backpropagation computes the gradient for every weight; gradient descent, or a variant such as Adam, is the separate step that uses that gradient to actually update the weights.

When should I use backpropagation instead of a gradient-free method?

Whenever the network is built from differentiable operations, which almost all standard neural networks are. Gradient-free methods are reserved for the discrete, non-differentiable steps backpropagation cannot handle, such as selecting one of several actions.

What are the risks or limits of backpropagation?

Vanishing or exploding gradients in very deep networks, no support for genuinely non-differentiable operations, and a memory cost that scales with every parameter in the model.

Does backpropagation require GPUs?

Not strictly, but training anything beyond a small network is impractically slow on a CPU, because the backward pass repeats the same matrix operations as the forward pass across every layer and every training example.

How much does backpropagation cost to run?

Roughly two to three times the compute of a forward pass alone, plus extra GPU memory to hold the intermediate values the backward pass needs, scaling with the number of parameters in the network.

Who actually invented backpropagation?

Seppo Linnainmaa described the underlying method in a 1970 thesis, and Paul Werbos applied a related idea to neural networks in 1974 and 1982, years before Rumelhart, Hinton and Williams's 1986 paper made it well known.

Keep reading

More from AI

All of AI
All of AI