011 min
Loss function at a glance
- What it is: a single number that scores how wrong one prediction is.
- Why it exists: training needs a number that gets smaller as the model improves.
- What it costs: the wrong choice trains a model toward the wrong goal, silently.
- When it breaks: a falling loss number does not guarantee the model is actually improving.
- As of: cross-entropy for classification and mean squared error for regression are still the defaults.
022 min
The problem a loss function solves
Before a model can be trained, "getting better" has to become something a computer can act on. Early approaches to automatic learning ran into this directly: counting how many predictions were right or wrong tells a person whether a model is good, but it gives a training procedure nothing to work with, because flipping one prediction from wrong to right is a jump with no steps in between.
Rumelhart, Hinton, and Williams, describing the backpropagation procedure that made training multi-layer networks practical, framed the entire learning problem around one number instead. The procedure they described repeatedly adjusts the weights of the connections in a network to minimize a measure of the difference between the network's actual output and the output it should have produced. That measure is the loss function. Every weight in the network gets changed, over and over, in whichever direction makes that one number smaller.
This sets a real requirement on what the number can be. It has to change smoothly as a weight changes by a tiny amount, because the training procedure needs to know which direction, out of millions of possible small adjustments, actually reduces the number. Plain accuracy, the fraction of predictions that were exactly right, does not change smoothly like this: it stays flat while a wrong prediction gets slightly less wrong, then jumps the moment the prediction crosses over to correct. A goal that cannot be expressed this way cannot be the thing a model is directly trained to minimize, whatever a person actually cares about.
032 min
How a loss function works
A loss function takes two things for one training example β what the model predicted, and what the correct answer was β and returns a single number that grows larger the more wrong the prediction was. Training repeats one cycle: compute this number across a batch of examples, work out how each weight in the model would need to change to make the number smaller, and change every weight slightly in that direction. That weight-changing step is Backpropagation; the loss function is what tells backpropagation which direction reduces the number.
Take a spam filter deciding whether one email is spam or not. For each email, the model outputs a probability for each of the two categories β say it assigns 70% to "spam" and 30% to "not spam" for an email that is actually spam. The standard loss for this kind of prediction is cross-entropy. Stanford's CS231n course notes describe what the softmax version of this loss is actually doing: once the model's output is read as a probability, minimizing cross-entropy means minimizing the negative log likelihood of the correct class. In practice, that means the loss on this one email depends only on the 70% the model assigned to "spam," the correct category β not on how the remaining 30% was split, because with only two categories there is nowhere else for it to go. The name comes from information theory, where cross-entropy measures the gap between two probability distributions; for one labeled example, the "true" distribution puts all of its weight on the one correct category.
Cross-entropy loss for one example: L = βlog(p), where p is the probability the model assigned to the correct category.
Regression needs a different shape of loss, because there is no fixed set of categories to assign a probability to. Predicting a number β a delivery estimate in days, a price, a temperature β uses squared error instead: the loss is the square of how far the prediction was from the correct value. Google's Machine Learning Crash Course explains the effect of that squaring plainly: a large miss counts for much more than a small one, and a small miss counts for even less, because the difference is squared rather than measured directly.
Both shapes of loss meet the same requirement backpropagation needs: the number has to change smoothly as a weight changes by a tiny amount, with no abrupt jumps. A function built around "exactly right or not," the way plain accuracy is scored, cannot meet that requirement. That is why accuracy is essentially never the number a model is trained on directly, even though it is usually the number people actually care about.
041 min
AlexNet and the 2012 ImageNet benchmark
In 2012, Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton trained an eight-layer convolutional neural network, later known as AlexNet, on the ImageNet dataset's 1,000-category classification task. The paper describing it states the loss the network was actually trained on plainly: the network maximizes the multinomial logistic regression objective, which is equivalent to maximizing the average across training cases of the log-probability of the correct label under the prediction distribution. That is the same quantity as minimizing cross-entropy loss, stated the other way around β maximizing the log-probability of the right answer and minimizing the negative log-probability of the right answer are the same optimization problem.
Nothing about "recognize objects correctly" was written into the network directly. The network only ever saw this one number, computed fresh for every training image, and adjusted its weights, through backpropagation, to make that number smaller, millions of times over. Entered in the 2012 ImageNet Large Scale Visual Recognition Challenge, the network achieved a winning top-5 test error rate of 15.3%, compared to 26.2% achieved by the second-best entry. A gap of that size, on a competition the whole field was watching, is widely credited with convincing researchers that deep networks trained this way were worth taking seriously β a result produced entirely by driving one loss number down, with no separate instruction about what "correct" should look like.
052 min
A worked example: computing loss by hand
A numeric example, worked by hand, shows the same mechanism from the inside. Suppose an image classifier is shown one photo of a cat, and cross-entropy loss is computed for four different levels of confidence the model might have assigned to the correct label, "cat":
| Probability assigned to the correct label | Cross-entropy loss (βlog p) |
|---|---|
| 0.99 | 0.01 |
| 0.70 | 0.36 |
| 0.30 | 1.20 |
| 0.01 | 4.61 |
The loss barely moves between a confident, correct 99% and a merely correct 70%. It grows sharply between an unsure 30% and a nearly-wrong 1%, because βlog(p) rises toward infinity as p falls toward zero. A model that assigns 1% to the right answer pays a loss more than 450 times larger than a model that assigns 99%, even though a system that only checks the top prediction would call both of them "right."
Squared error, used for regression, behaves differently because it has no probability to work from, only a plain numeric gap. Picture a model forecasting a package's delivery time in days, where the true answer is 3 days. A forecast of 4 days, off by one, contributes a squared error of 1Β² = 1. A forecast of 7 days, off by four, contributes 4Β² = 16 β sixteen times as much, not four times as much. Squared error does not care whether the model was confident; it cares only how far off the number was, amplified by the square. That is the one variable that separates the two shapes: cross-entropy reacts to how sure the model was, squared error reacts only to how far off the number was.
062 min
Loss function vs. nearby concepts
Loss function gets confused with three nearby terms, and each is worth separating out.
An objective function is the broader term: whatever quantity training is set up to optimize, whether that means minimizing a loss or maximizing something else, such as expected reward in reinforcement learning. Every loss function is an objective function pointed toward minimizing; not every objective function is a loss.
An evaluation metric is what a person actually reads to decide whether a model is good enough β accuracy, precision on one category, a business-specific error rate. It only has to be meaningful to a person; a loss function only has to be smooth enough for Backpropagation to compute a direction from. That difference is why a classifier is trained on cross-entropy but judged on accuracy: accuracy has no smooth gradient to follow, so it cannot do the training's job, even though it is what the training is ultimately for.
A regularization term gets added into the same final number a model minimizes, which is where the confusion with loss comes from, but it does a different job. Loss measures how wrong one prediction was; a regularization term penalizes the model for being unnecessarily complex, independent of whether any prediction was right, specifically to discourage the kind of memorizing that leads to Overfitting.
The common loss functions themselves also get blurred together:
| Loss | Used for | Behavior | Typical use |
|---|---|---|---|
| Cross-entropy | Classification | Punishes a confident wrong answer sharply; barely penalizes a slightly-unsure correct one | Spam filters, image classifiers, language models |
| Mean squared error (MSE) | Regression | Squares the gap, so a large miss counts far more than a small one | Price or demand forecasting |
| Mean absolute error (MAE) | Regression | Grows in a straight line with the size of the gap, so one outlier does not dominate training | Forecasts where a single bad outlier should not skew the model |
Choosing between MSE and MAE is a decision about outliers: MSE is the right call when large misses matter more than small ones; MAE is the right call when one extreme mistake, a supply-chain shock on an otherwise ordinary forecast, should not be allowed to dominate everything the model learns from every other day.
072 min
What choosing a loss function means for your work
For an engineer setting up a new model, the loss function is not a configuration detail decided after the fact β it is the actual definition of what the model is being trained to do, chosen before training starts. Picking cross-entropy for a task that is really about numeric closeness, or squared error for a task that is really about category membership, trains a working model toward the wrong target, and nothing in the training run itself flags this as a mistake; the loss still goes down.
For a PM reading a report that "training loss is going down" as evidence of progress, the honest reading is narrower than it sounds: it shows the model is getting better at the specific number it was told to minimize, not necessarily at the outcome the product needs. A support-ticket classifier can drive its cross-entropy loss steadily lower while still misrouting the rare, high-value ticket category a business cares about most, because the loss function weighs every example by how often it appears in the training data, not by how much a mistake on it costs.
For a founder or product lead deciding what "done" means for a model before launch, the loss function answers a narrower question than the one that matters commercially. A fraud model with 1 real case in 10,000 transactions can push cross-entropy very low by predicting "not fraud" almost every time, since the loss has no way to know which mistakes are the expensive ones β which is why a separate, human-chosen number, like recall on the fraud category specifically, has to be checked before deciding a model is ready, rather than trusting the loss number on its own.
081 min
What a loss function does not guarantee
A falling loss number does not guarantee the model is getting better at the task a person actually cares about. If the loss function is a poor stand-in for the real goal β squared error rewarding a model for shaving a little off every prediction rather than getting the rare, extreme case right β the model optimizes exactly what it was told to, and that turns out not to be the same thing as useful.
A falling training loss also does not mean the model generalizes to new data. A model can keep driving its loss on the training set toward zero long after it has stopped learning anything transferable, by memorizing the specific training examples instead β the failure named Overfitting β which is why training loss is always checked against loss on data the model never trained on, not read on its own.
A loss function can also look healthy for the wrong reason if the training data itself is compromised. When information about the correct answer leaks into what the model is allowed to see during training, a problem called Data Leakage, the loss falls convincingly during training and validation alike, while teaching the model nothing it can use once that leaked information is gone at prediction time.
091 min
What computing a loss function costs
The arithmetic of a loss function is cheap on its own β a handful of operations per training example. The real cost is that the loss has to be computed, and then differentiated back through every weight in the model, for every example, every time the weights update. For a large model trained on a large dataset, that adds up to an enormous number of these calculations across one training run, which is why training cost scales with model size and dataset size together, not with the loss formula itself.
There is a second, quieter cost. A mismatched loss function is rarely caught by any automated check, because the number still goes down; nothing in a training log distinguishes a model learning the right thing from one learning the wrong thing efficiently. Catching the mismatch usually takes a person checking a separate, human-meaningful metric against the loss, not the training run on its own.
?5 questions
Questions people ask
What is a loss function in simple terms?
What is an example of a loss function?
Is a loss function the same as accuracy?
What is the difference between a loss function and a cost function?
Does a low loss mean a model is good?
Β§4 sources
Sources
Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323, 533β536.
Karpathy, A. Linear Classification: Support Vector Machine, Softmax. CS231n Course Notes, Stanford University.
Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet Classification with Deep Convolutional Neural Networks. Advances in Neural Information Processing Systems 25.
Google Machine Learning Crash Course. Linear Regression: Loss.





