011 min
Overfitting at a Glance
- What it is β a model that memorizes its training data instead of learning the pattern behind it.
- Why it exists β models with enough capacity will fit noise if nothing stops them.
- What it costs β confident, wrong predictions on exactly the data the model will actually see in production.
- When it breaks β training accuracy keeps climbing while validation accuracy stalls or falls.
- As of 2026 β still the default failure mode diagnosed with a train/validation loss curve, not a new problem specific to large models.
021 min
The Problem Overfitting Solves
Before overfitting had a name, the naive approach to fitting a model was to make it match the training examples as closely as possible and call the job done. That approach fails the moment a model is expressive enough to fit not just the real pattern in the data, but also the noise: the specific, unrepeatable quirks of exactly the examples it happened to see. A model that memorizes noise gets every training example right and then gets new examples wrong, because the noise in a new batch of data is different noise.
The failure is specifically about a model that looks like it is succeeding. Training accuracy climbs, the loss curve drops, and every internal metric available during training says the job is going well. The only way to catch overfitting is to check performance on data the model never trained on, which is why a validation set exists as a separate concept from the training set at all.
031 min
How Overfitting Works
Picture a model learning to classify emails as spam. If the training set happens to contain twelve spam emails that all mention a specific promotional phrase, an overfit model can learn "contains this exact phrase" as a rule, rather than learning the broader pattern of what spam actually looks like. On the training set, that narrow rule scores perfectly, because every spam example the model saw happened to contain the phrase. On a new batch of email, where spam uses different wording, the narrow rule fails silently.
The standard way to catch this is watching two numbers during training instead of one: the loss on the training set, and the loss on a validation set the model never trains on. Training loss and validation loss track each other closely at first, and then diverge: training loss keeps falling while validation loss levels off or climbs. That divergence point is the moment the model stops learning the general pattern and starts memorizing the specific training examples instead.
A model's capacity to overfit tracks its capacity in general: more parameters, more training passes, and less training data all give a model more room to fit noise rather than signal. This is the same underlying issue whether the model is a three-parameter polynomial or a billion-parameter neural network β only the scale of what gets memorized changes.
041 min
A Concrete Example
Google Flu Trends is the clearest documented case of overfitting happening at a massive scale outside a textbook. Launched to estimate flu prevalence from search-query patterns, the system's initial methodology, according to a 2014 analysis in Science by David Lazer, Ryan Kennedy, Gary King, and Alessandro Vespignani, was to find the best matches among 50 million search terms to fit 1152 data points. With fifty million candidate variables and just over a thousand real data points to fit them against, the odds of finding search terms that happened to correlate with flu-like symptoms for reasons that had nothing to do with actual flu were high. The system had missed high for 100 out of 108 weeks starting with August 2011 by the time researchers examined it closely β a model that had learned coincidental search patterns in its original training window rather than anything causally tied to flu.
051 min
What This Means for Your Work
For a PM, overfitting shows up as a model that scores brilliantly in a demo built from the same data pipeline used to train it, and then embarrasses the team the week after launch on real user traffic the training set never saw. The decision this changes is what to trust as evidence a model is ready to ship: a validation score on data the model has genuinely never touched, never a training score, and never a demo built by hand-picking flattering examples.
For an engineer building an eval pipeline, overfitting shows up as a model or prompt that has been iterated against the same small set of test cases so many times that it has effectively memorized the answers to that specific set, the same failure mode as a model overfitting training data, just happening to the evaluation process itself. This is why an eval set needs adversarial cases and a held-out slice that nobody optimizes against directly, not just a growing pile of happy-path examples everyone keeps re-running.
For a designer running user tests, the analogous mistake is treating results from the same five repeat testers as representative of the whole user base β a human version of training and validating on the same data, where confidence keeps climbing for reasons that have nothing to do with whether the design actually works for anyone else.
061 min
What Overfitting Costs
The direct cost of overfitting is not compute or latency; it is a false sense of certainty. A team that ships based on training or demo metrics inherits every gap between the training distribution and the real world, discovered only after the model is live and making real decisions. Diagnosing it costs a held-out validation set that is never touched during model development, which in a data-scarce project can mean giving up ten to twenty percent of the available labeled data purely to have a trustworthy number to check against.
The Google Flu Trends failure carried a public cost too: the system had been "persistently overestimating flu prevalence for a much longer time" than was widely reported before the 2014 analysis caught it, meaning the errors it introduced into public-health monitoring ran for years before anyone outside the project traced them back to how the model had been fit.
Running that check is not free either: a team has to actually wait for a validation score before trusting a model, rather than shipping the moment training loss looks good, which is a real discipline cost on top of the data cost.
071 min
What Overfitting Does Not Solve
Overfitting is detected, not prevented, by watching where training and validation performance diverge. The underlying cause is usually one of two things: the training set doesn't fairly represent the data the model will actually see, or the model has more capacity than the amount of training data can responsibly support. Both causes point to the same two families of fix.
The first family reduces what the model is allowed to memorize: regularization techniques that penalize complexity directly, early stopping that halts training at the point the validation loss starts climbing rather than training until the training loss bottoms out, and dropout or similar techniques that prevent any single piece of the model from depending too heavily on any single training example.
The second family gives the model less noise to memorize in the first place: more training data, more diverse training data, and removing near-duplicate examples that let the model "memorize" a pattern by brute repetition rather than by generalizing. Cross-validation, running the same train/validate split several different ways and averaging the results, catches the case where a single validation split happened to be unusually easy or unusually hard, and so looked fine even though the underlying model was overfit.
081 min
Overfitting vs. Nearby Concepts
The nearest neighbor is underfitting, and the two sit on opposite ends of the same axis: an overfit model is too complex for its data and has memorized noise; an underfit model is too simple for its data and has failed to learn even the real pattern. Both produce a model that performs badly on new data, but an underfit model also performs badly on the training data it was given, while an overfit model looks excellent on training data and only reveals the problem on a validation set. The deciding fact is which loss curve is high: if training loss itself is high, the model is underfit; if training loss is low but validation loss is high, it is overfit.
A second neighbor is data leakage, which produces the same symptom, a suspiciously good validation score, through a different cause: leakage happens when information from outside the legitimate training data β a duplicate row that ended up in both the training and validation split, or a feature that encodes the answer β sneaks into training. Overfitting is a capacity problem; leakage is a data-hygiene problem, and no amount of regularization fixes leakage, because the model isn't wrong to trust information that is genuinely, if illegitimately, available to it.
091 min
A Second Case
A second, very differently shaped case of overfitting comes from medical imaging rather than search data. John Zech and colleagues trained pneumonia-screening convolutional neural networks on 158,323 chest x-rays from three hospital systems and found that in three out of five natural comparisons, performance on chest x-rays from outside hospitals was significantly lower than on held-out x-rays from the original hospital systems. Investigating why, they found the CNNs were able to detect where an x-ray was acquired, meaning which hospital system or hospital department, with extremely high accuracy and calibrate predictions accordingly.
Where Google Flu Trends overfit to coincidental search terms across the whole country, the radiology models overfit to something narrower and stranger: equipment and labeling differences between hospitals that correlated with disease prevalence at each specific hospital, but had nothing to do with disease itself. A model that looks like it is reading an x-ray for signs of pneumonia can, without anyone intending it, actually be reading which hospital's scanner produced the image.
101 min
Where the Evidence Is Contested
Not every gap between training and validation performance is evidence of overfitting in the strict sense, and treating every such gap the same way is itself a common mistake in how teams diagnose model problems. A validation set drawn from a different time period, user population, or data source than the training set will show a performance gap even for a perfectly regularized model, because the two sets are not really measuring the same distribution β the Zech radiology study is itself an example of this ambiguity: the paper's own summary notes the performance gap reflects not only the model's ability to identify disease-specific imaging findings, but also its ability to exploit confounding information, meaning some of what looks like "overfitting" is really a model correctly learning patterns in its training hospitals that simply do not hold at other hospitals.
This distinction matters in practice: regularizing a model harder does not fix a validation set that measures the wrong thing, and teams that respond to every train/validation gap with more dropout or more data, without first checking whether the two sets are actually comparable, can spend real engineering effort solving the wrong problem.
?6 questions
Questions people ask
What causes overfitting?
What is an example of overfitting in a real product?
How do you avoid or fix overfitting?
What is the difference between overfitting and underfitting?
Does overfitting require a huge model?
How much does checking for overfitting cost?
Β§4 sources
Sources
Overfitting β Google Machine Learning Crash Course.
Lazer, D., Kennedy, R., King, G., & Vespignani, A. (2014). The Parable of Google Flu: Traps in Big Data Analysis. Science, 343(6176), 1203β1205.
Zech, J. R., Badgeley, M. A., Liu, M., Costa, A. B., Titano, J. J., & Oermann, E. K. (2018). Confounding Variables Can Degrade Generalization Performance of Radiological Deep Learning Models. arXiv:1807.00431.
Interpreting Loss Curves β Google Machine Learning Crash Course.





