Activation Function

An activation function is the rule a single neuron in a neural network applies to its combined input before passing a signal onward, turning a plain weighted sum into a nonlinear value so the network can learn curves rather than only straight lines.

12 min read

By Ravi SuranaUpdated 7 sources

Quick answer

~20 sec

An activation function is the rule a single neuron in a neural network applies to its combined input before passing a signal onward. It turns a plain weighted sum into a nonlinear value, often zero for anything negative and the same number for anything positive, so a stack of neurons can learn curves rather than only straight lines.

011 min

Activation Function at a glance

  • What it is: A fixed nonlinear rule a neuron applies to its weighted-sum output.
  • Why it exists: Stacked linear layers collapse into one straight line without it.
  • What it costs: GELU and Swish cost more compute per neuron than ReLU.
  • When it breaks: A ReLU neuron can permanently stop firing, called the dying ReLU problem.
  • As of: ReLU and GELU remain the two most common defaults in 2026 production models.

021 min

The problem Activation Function solves

Before activation functions were added between layers, the natural way to build a bigger network was to stack more layers, each one only computing a weighted sum of the layer before it. That approach fails for a precise mathematical reason: a weighted sum of a weighted sum is still just a weighted sum. Stack a hundred such layers and the whole network still computes nothing more than one giant straight-line combination of the input, no matter how deep it goes.

Real patterns are rarely straight lines. A model trying to decide whether a photo shows a dog cannot use a rule like "brighter pixels always mean more dog" β€” past a certain point, brighter just means overexposed, and the relationship bends the other way. A network built only from stacked weighted sums cannot represent that bend, however many layers it has; it can only rotate and rescale its input, never reshape it. Every useful deep network needs at least one place in the pipeline where the shape of a relationship, not just its scale, can change. That is the job an activation function does, at every single neuron, on every layer.

032 min

How Activation Function works

Picture a single neuron near the end of a phone's photo-tagging model, the part deciding whether one photo should be tagged "dog." That neuron does not look at the whole photo. It receives a handful of numbers from the layer before it β€” one standing for how much the fur texture looks right, one for the ear shape, one for the snout shape β€” multiplies each by a weight the network learned during training, and adds them up. Call that sum z. A photo with strong fur, ear and snout evidence might produce z = 6. A photo of a cat, with weak or negative evidence on all three, might produce z = -4.

On its own, z is just a number, and a chain of such sums, however many layers deep, only ever computes one more weighted sum of the original pixels β€” the reason a network needs more than stacked sums at all is covered above. The activation function is the step, applied to z at every single neuron, that turns that raw sum into something the next layer can use differently depending on its size.

ReLU(z) = z when z is positive, and 0 when z is negative or zero.

The simplest and still most common choice, the rectified linear unit or ReLU, does exactly that: z = 6 passes through unchanged as 6, and z = -4 becomes 0. The neuron with strong dog evidence keeps contributing its full strength to the next layer; the neuron with negative evidence contributes nothing at all, as if it had never fired. That single rule, pass it through or block it entirely, is what lets a network build sharp, threshold-like decisions instead of only smooth blends.

A different activation function, the logistic sigmoid, squashes z into a fixed range between 0 and 1 instead of gating it at zero: a large positive z comes out close to 1, a large negative z comes out close to 0, and the values in between get compressed into an S-shaped curve. Squashing like this earns its keep exactly where the next step needs something that reads like a probability β€” how confident the network is that the photo shows a dog β€” rather than a raw, unbounded score.

The step where the interesting thing happens is always the same: a neuron's incoming evidence, however it was gathered, gets reshaped by a fixed nonlinear rule before it can affect anything downstream. Change that rule and the same z produces a different signal, which is why the choice of activation function is a real design decision and not a default nobody has to think about.

041 min

A concrete example

In 2013, Andrew Maas, Awni Hannun and Andrew Ng at Stanford tested this choice directly instead of assuming a nonlinear function's exact shape does not matter. They trained the same deep network architecture two ways β€” once with the sigmoidal nonlinearity standard at the time, once with ReLU β€” as an acoustic model for the 300-hour Switchboard conversational speech recognition task: turning raw audio into the words a speaker said. Using simple training procedures without pretraining, networks with rectifier nonlinearities produce 2% absolute reductions in word error rates over their sigmoidal counterparts. On a task where every fraction of a percentage point is fought over, a 2-point absolute drop in word error rate, from swapping one function for another inside every neuron, is a large result.

The gain held up, and grew, as the networks got deeper. Each additional hidden layer helped the ReLU networks more than it helped the sigmoidal ones, because the sigmoidal networks were quietly losing their training signal in their lower layers β€” the failure covered below β€” while the ReLU networks were not. The experiment is a clean before-and-after: same audio, same architecture, same training procedure, only the activation function changed, and the error rate moved by two full points because of it.

051 min

What this means for your work

An engineer picking an architecture for a model that will run on a phone, not a data center, treats the activation function as a real compute cost, not a free stylistic choice. Multiplied across millions of neurons and thousands of requests a day, a heavier activation function shows up directly in battery drain and response time, so whatever smoothness it buys has to be worth what it costs β€” the cost section below spells out exactly why.

A PM reading a vendor's benchmark claims needs to ask what changed between the old model and the new one before trusting a headline accuracy number. Swapping ReLU for GELU or Swish alone typically moves accuracy by well under a single percentage point, so a benchmark advertising a much larger jump almost always changed something bigger than its activation function, and a PM who does not ask will credit the wrong improvement.

A founder deciding whether to fine-tune an existing model or train a smaller one from scratch runs into the same choice one layer up: an older or smaller open-source model may still default to plain ReLU because it predates GELU's adoption, and matching that architecture's original activation function, not guessing a newer one, is usually what makes a from-scratch retrain match the vendor's published numbers at all.

061 min

What Activation Function costs

ReLU's cost is close to nothing: one comparison against zero, run at every neuron, on every forward pass through the network. That cheapness is a real reason it stayed the default for over a decade even after smoother alternatives existed.

GELU and Swish cost more. Computing GELU exactly needs the Gaussian error function; most implementations approximate it, but the approximation still needs a multiply and an exponential-like term at every neuron, not just one comparison. Swish needs a full logistic-sigmoid evaluation at every neuron for the same reason. In a small network nobody would notice the difference. In a model with billions of parameters, run millions of times a day, that per-neuron cost is multiplied by the model's entire size and every request it serves, which is why the choice is a genuine ongoing compute cost and not only a training-time detail.

The gradient computed during backpropagation also passes back through the activation function at every neuron, so a more expensive activation function costs more on the way back through the network during training, not only on the way forward at inference.

072 min

What Activation Function does not solve

A ReLU neuron can get permanently stuck. If its weights end up producing a negative z for every example in the training set, its output is 0 every time and its gradient is 0 every time, and nothing in ordinary gradient-based training can revive it. Researchers Maas, Hannun and Ng described the exact mechanism: the gradient is 0 whenever the unit is not active. This could lead to cases where a unit never activates as a gradient-based optimization algorithm will not adjust the weights of a unit that never activates initially. A network can end up with a meaningful fraction of its neurons dead this way, quietly wasting capacity the model paid for but can no longer use. Leaky ReLU, which lets a small negative slope through instead of a hard zero, exists specifically to keep that gradient alive so a stuck neuron can recover.

Sigmoid and tanh fail differently, through saturation rather than death. Maas and colleagues define the mechanism plainly: vanishing gradients occur when lower layers of a DNN have gradients of nearly 0 because higher layer units are nearly saturated at -1 or 1, the asymptotes of the tanh function. A deep sigmoid network can end up training so slowly its lower layers barely change at all.

Neither failure is overfitting, and swapping activation functions does not fix overfitting β€” that comes from a model with too much capacity for too little data, not from the shape of its nonlinearity.

081 min

Activation Function vs. nearby concepts

The concept most often confused with an activation function is the loss function. Both sit inside the same forward-and-backward training loop, and both get called "a function" in almost every diagram, which is most of why beginners mix them up.

The deciding fact is where each one sits and how many times it runs. An activation function runs once per neuron, at every layer, on every single forward pass β€” it is part of what the network computes to produce an answer at all. A loss function runs once per training example, only at the very end, after the network has already produced its output β€” it compares that output to the correct answer and produces the one number the network tries to shrink.

Activation functionLoss function
RunsAt every neuron, every layerOnce, at the output
JobShapes one neuron's signalScores the whole prediction
Used at inferenceYesNo

A network needs both to train at all, but only one of them, the activation function, is still doing anything once training stops and the model is just answering questions.

092 min

How Activation Function changed since

Early neural networks, including the 1958 perceptron, used a hard step function: fire or do not fire, with nothing in between. Sigmoid and tanh replaced the step function because they are smooth enough to train with gradient-based methods, and they were the default through the 1990s.

Rectified linear units changed that default. Vinod Nair and Geoffrey Hinton's 2010 paper brought the rectifier into deep learning. Unlike binary units, rectified linear units preserve information about relative intensities as information travels through multiple layers of feature detectors. Xavier Glorot, Antoine Bordes and Yoshua Bengio's 2011 study went further, creating sparse representations with true zeros, which seem remarkably suitable for naturally sparse data β€” meaning most neurons in a trained ReLU network sit at exactly zero for any given input, not just near it. In 2012, Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton's AlexNet made the practical case impossible to ignore. Their four-layer convolutional network reaches a 25% training error rate on CIFAR-10 six times faster than an equivalent network with tanh neurons, and that speed difference is part of what let them train a network large enough to win the 2012 ImageNet competition at all.

ReLU has not been replaced since, but it has been joined. Dan Hendrycks and Kevin Gimpel introduced a smoother alternative in 2016: the GELU nonlinearity weights inputs by their value, rather than gates inputs by their sign as in ReLUs. Prajit Ramachandran, Barret Zoph and Quoc Le found a second smooth alternative, Swish, in 2017, using automated search rather than hand design. Neither replaced ReLU outright; both are now standard options alongside it, chosen per architecture rather than as a universal upgrade.

101 min

A second case

The Switchboard speech-recognition experiment above shows an activation function doing its most basic job: gating a hidden neuron's signal on or off. A different, later case shows the same concept doing a different job inside a much bigger system.

By 2018, transformer-based language models were replacing ReLU inside their hidden layers with the smoother GELU. BERT's own paper explains the choice directly: its authors used GELU rather than the standard relu, following OpenAI GPT. The same architecture also uses a completely different activation function, softmax, in its attention layers and its final output layer, turning a set of raw scores into a probability distribution that sums to exactly 1 across every candidate word.

The variable that changes between the two cases is what the activation function is for, not just which one gets used. In the 2013 speech model, the job was gating: let strong evidence through, block weak evidence entirely. In a transformer's hidden layers, GELU still gates, but more smoothly, because a hard cutoff at zero cost a small but real amount of accuracy at that scale. In the same transformer's output layer, softmax gates nothing β€” it reshapes a set of numbers into something that can be read as a probability. Activation function is not one interchangeable part; it is a category whose right shape depends on the job a given layer is doing.

111 min

Where the evidence on Activation Function is contested

Not everyone agrees the specific choice of activation function matters as much as the papers proposing new ones suggest. Prajit Ramachandran, Barret Zoph and Quoc Le, the researchers who discovered Swish by searching automatically rather than designing it by hand, said so themselves in the same paper that introduced it: although various hand-designed alternatives to ReLU have been proposed, none have managed to replace it due to inconsistent gains.

Their own numbers show why the doubt persists. In their best result, simply replacing ReLUs with Swish units improves top-1 classification accuracy on ImageNet by 0.9% for Mobile NASNet-A and 0.6% for Inception-ResNet-v2 β€” real, measured gains, but small enough that many teams judge the extra compute cost from the earlier section not worth paying, and stick with plain ReLU rather than chase a fraction of a point.

The more contested question is whether these differences generalize. A gain measured on one architecture and one dataset does not reliably predict the gain on a different one, which is why practitioners who have watched several "ReLU replacements" arrive with strong claimed results treat each new one the same way: worth testing on the actual model in question, not worth adopting on the strength of somebody else's benchmark alone.

122 min

Frequently asked questions about Activation Function

What is an activation function?

An activation function is the nonlinear rule a neural network applies to a single neuron's weighted-sum input before passing a signal to the next layer. It is what lets a network learn curved relationships instead of only straight lines.

What is the difference between ReLU and sigmoid?

ReLU passes positive values through unchanged and blocks negative values to zero. Sigmoid squashes every value into a fixed range between 0 and 1, which is useful when the output needs to read like a probability rather than a raw score.

Why do neural networks need an activation function at all?

Without one, stacking any number of layers is mathematically identical to a single layer, because a weighted sum of weighted sums is still just a weighted sum. Only a nonlinear step lets depth add real learning power.

What is the dying ReLU problem?

A ReLU neuron whose weights always produce a negative input outputs zero for every example and has a gradient of zero, so ordinary training can never revive it. The neuron is permanently inactive.

Is GELU better than ReLU?

GELU tends to score slightly higher on many benchmarks and is standard in large language models like BERT and GPT, but it costs more compute per neuron than ReLU, so the trade-off is not automatic.

Does an activation function require a GPU or extra training data?

No. It adds a small, fixed amount of compute at every neuron on every pass, but it needs no separate training data or hardware beyond whatever already runs the network itself.

How much does an activation function cost to run?

ReLU costs close to nothing, about one comparison per neuron. GELU and Swish cost more, typically an extra multiply and an exponential-like term per neuron, which adds up across a model with billions of neurons.

?7 questions

Questions people ask

What is an activation function?

An activation function is the nonlinear rule a neural network applies to a single neuron's weighted-sum input before passing a signal to the next layer. It is what lets a network learn curved relationships instead of only straight lines.

What is the difference between ReLU and sigmoid?

ReLU passes positive values through unchanged and blocks negative values to zero. Sigmoid squashes every value into a fixed range between 0 and 1, which is useful when the output needs to read like a probability rather than a raw score.

Why do neural networks need an activation function at all?

Without one, stacking any number of layers is mathematically identical to a single layer, because a weighted sum of weighted sums is still just a weighted sum. Only a nonlinear step lets depth add real learning power.

What is the dying ReLU problem?

A ReLU neuron whose weights always produce a negative input outputs zero for every example and has a gradient of zero, so ordinary training can never revive it. The neuron is permanently inactive.

Is GELU better than ReLU?

GELU tends to score slightly higher on many benchmarks and is standard in large language models like BERT and GPT, but it costs more compute per neuron than ReLU, so the trade-off is not automatic.

Does an activation function require a GPU or extra training data?

No. It adds a small, fixed amount of compute at every neuron on every pass, but it needs no separate training data or hardware beyond whatever already runs the network itself.

How much does an activation function cost to run?

ReLU costs close to nothing, about one comparison per neuron. GELU and Swish cost more, typically an extra multiply and an exponential-like term per neuron, which adds up across a model with billions of neurons.

β–Ά3 videos

Watch

  • Understanding Activation Function

  • Activation Functions In Neural Networks Explained | Deep Learning Tutorial

    AssemblyAI

  • Activation Functions | Deep Learning Tutorial 8 (Tensorflow Tutorial, Keras & Python)

    codebasics

Β§7 sources

Sources

  1. Nair, V., & Hinton, G. E. (2010). Rectified Linear Units Improve Restricted Boltzmann Machines. ICML 2010.

  2. Glorot, X., Bordes, A., & Bengio, Y. (2011). Deep Sparse Rectifier Neural Networks. AISTATS 2011.

  3. Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet Classification with Deep Convolutional Neural Networks. NeurIPS 2012.

  4. Maas, A. L., Hannun, A. Y., & Ng, A. Y. (2013). Rectifier Nonlinearities Improve Neural Network Acoustic Models. ICML Workshop on Deep Learning for Audio, Speech and Language Processing.

Show all 7 sources
  1. Hendrycks, D., & Gimpel, K. (2016). Gaussian Error Linear Units (GELUs). arXiv:1606.08415.

  2. Ramachandran, P., Zoph, B., & Le, Q. V. (2017). Searching for Activation Functions. arXiv:1710.05941.

  3. Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805.

Keep reading

More from AI

All of AI
All of AI