011 min
Online Learning at a Glance
- What it is β updating a model one example at a time as data streams in, instead of retraining on a fixed batch.
- Why it exists β a model trained once goes stale the moment new data arrives faster than a full retrain can absorb it.
- What it costs β a live feedback pipeline and careful tuning; a model that reacts to every point can also overreact to noise.
- When it breaks β when incoming data is missing, mislabeled, or arrives with no reliable feedback signal at all.
- As of 2026 β large-scale online logistic regression, first published by Google's ad-prediction team in 2013, still runs in production.
022 min
The Problem Online Learning Solves
Before online learning was common, the standard way to train a predictive model was batch learning: collect a fixed set of training examples, train until the model converges, then freeze it and serve it. That works well when the world the model describes stays still. It fails the moment new data keeps arriving faster than a full retrain can absorb it β a fraud detector, a news feed ranker, or an ad system updated once a day is, for the rest of that day, working off yesterday's picture of its users.
Steven Hoi and colleagues, in a widely cited 2018 survey of the field, name the resulting failures plainly: "Traditional batch learning methods suffer from some critical drawbacks: (i) low efficiency in both time and space costs; and (ii) poor scalability for large-scale applications because the model often has to be re-trained from scratch for new training data." Retraining from scratch is expensive, and by the time the new model ships, the data it trained on is already hours or days old.
Google's ad click-prediction team, describing a production system in a 2013 paper, put the same pressure in concrete terms: "It is necessary to make predictions many billions of times per day and to quickly update the model as new clicks and non-clicks are observed." Waiting for the next scheduled retrain was not an option at that volume β the model had to keep learning while it was still serving predictions.
033 min
How Online Learning Works
Online learning follows a simple loop, repeated once for every new example instead of once for an entire dataset. Picture a spam filter deciding, email by email, whether each new message is spam β the same running example a 2018 survey of the field uses to explain the loop: "Consider spam email detection as a running example of online binary classification, where the learner answers every question in binary: yes or no."
On each round, the filter receives one message, represented as a set of features β the words it contains, the sender, and so on β and predicts spam or not spam. Only after that prediction does it learn the true answer, usually because a person marked the message. Hoi and colleagues describe the same two steps in general terms: "On each round, a learner receives a data instance, and then makes a prediction of the instance, e.g., classifying it into some predefined categories. After making the prediction, the learner receives the true answer about the instance from the environment as a feedback."
The step where the interesting thing happens comes next. The filter scores how wrong its guess was with a loss function, then nudges its weights toward the correct answer for that one message before moving to the next: "Based on the feedback, the learner can measure the loss suffered, depending on the difference between the prediction and the answer. Finally, the learner updates its prediction model by some strategy so as to improve predictive performance on future received instances."
The earliest algorithm built this way is the perceptron, introduced by Frank Rosenblatt in 1958. It updates its weights only when it gets a prediction wrong, moving them toward the correct label for that one example, then moving on. A 2018 survey still calls it the starting point: "Perceptron (Rosenblatt, 1958; Agmon, 1954; Novikoff, 1963) is the oldest algorithm for online learning."
Because an online learner never sees the whole dataset before it starts predicting, researchers judge it differently than a batch model. Martin Zinkevich, who formalized this in 2003, compares the learner's running loss to the best single fixed choice picked with hindsight: "We calculate our regret by comparing ourselves to an 'offline' algorithm that has all of the information before it has to make any decisions, but is more restricted than we are in the choices it can make." That gap has a name: "Regret is the difference between our cost and the cost of the offline algorithm." A good online learner is not one that never makes mistakes β it is one whose regret grows slowly as more examples arrive.
This also explains why online learning scales the way it does. Google's ad-prediction system trains continuously on a stream of ad impressions, and its engineers point to a specific reason for the shape of that update: "since each training example only needs to be considered once." A batch algorithm passes over the same fixed dataset many times until it converges; an online learner looks at each example once and moves on, because by the time it would want a second pass, newer examples have already arrived β as one of the paper's own footnotes puts it, "the name online emphasizes we are not solving a batch problem, but rather predicting on a sequence of examples that need not be IID."
041 min
A Concrete Example
Google's ad click-prediction system is one of the largest online learners running in production, updating continuously on the click and non-click stream described above. Its engineers published a direct, measured comparison of two ways to run the same online update rule on the same live data.
Before: the model used one single learning rate β the step size controlling how far each weight moves after one example β shared across every feature. A feature that appears in almost every ad and a feature that appears once a month were nudged by the same-sized step on every update.
After: each feature got its own learning rate, shrinking faster for features that update often and staying larger for features seen rarely β an approach the team calls per-coordinate learning rates.
The team reported the difference directly: "The results showed that using a per-coordinate learning rates reduced AucLoss by 11.2% compared to the global-learning-rate baseline." That is a large shift by the paper's own standard: "in our setting AucLoss reductions of 1% are considered large." The model, the data stream, and the update rule all stayed the same; only how fast each individual feature was allowed to adapt changed.
051 min
How Online Learning Changed Since
The perceptron's 1958 update rule was simple and worked, but it came with no general way to measure how good an online learner's guarantee was beyond counting mistakes on data that could be perfectly separated. Martin Zinkevich's 2003 paper reframed the whole problem in the language of convex optimization: instead of counting mistakes, measure how much worse the learner's running loss is than the best single fixed choice, picked with hindsight, over the same sequence. That reframing, called online convex programming, gave online learning the same mathematical footing as batch optimization, and it is the formulation most current online learning research still builds on.
The practical side moved just as far. What began as a small linear classifier for cleanly separable data is now, in the FTRL-Proximal system Google published a decade later, a logistic regression model with billions of coefficients updated continuously on a live production stream. The underlying idea β update from one example, then move on β has not changed since 1958. What changed is the scale it runs at and the theory used to say when it can be trusted.
063 min
What This Means for Your Work
For an engineer deciding how to ship a model that faces streaming data β a fraud detector, a feed ranker, a real-time bidding model β online learning changes what has to exist before launch. A batch pipeline needs a schedule and a place to store the next full training set; an online pipeline needs a live feature-and-label stream, a way to update the serving model without redeploying it, and a way to catch a bad update fast, because a bad batch of feedback now reaches production within minutes rather than waiting for a review before the next release.
When an online-updated fraud model's precision drops sharply within an hour of a new update, an engineer wants to see that drop immediately on a dashboard, so they can freeze further updates before the model learns from a batch of bad transactions.
That freeze-on-anomaly monitoring does not exist in most batch pipelines, because a batch model is reviewed once, before it ships, rather than while it is live.
For a PM deciding whether a feature needs this at all, the real question is not whether online learning is better but whether the data goes stale fast enough to justify the extra operational cost. A product catalog that changes once a week does not need it; a bidding system reacting to an auction running a thousand times a second does. Picking the wrong side of that line either wastes engineering effort keeping a slow-moving problem updated every second, or ships a model that is stale within the hour for a problem that needed the opposite.
For a founder scoping a first version of a personalization or fraud feature, online learning is often the more expensive path to build correctly, not the simpler one: it needs monitoring, a rollback path, and usually a fallback model to serve if the live stream stops updating β none of which a once-a-week batch retrain needs on day one. It is worth that cost mainly once the product has enough live traffic that stale predictions, rather than the model's raw quality, become the bigger source of lost revenue.
The thing this decision changes least is one many teams assume it changes most: the model's architecture rarely needs to differ between an online and a batch version of the same problem. The same logistic regression or gradient-boosted model can often be trained either way; what differs is the infrastructure and the operational discipline around it, which is why the build decision belongs to engineering and product together, not to model choice alone.
For a designer reviewing how a ranked feed or a set of recommendations behaves, online learning changes what consistency means day to day. A batch-trained feed looks the same all day, then shifts once at the next release; an online-updated feed can visibly shift within the same session a user is reading it, as their own clicks feed back into what they see next. That is often the point of the feature, but it also means a design review has to account for a feed that can look different an hour from now for reasons no single release note will explain.
Large deep learning models, including large language models, are the clearest case where a team can assume the wrong regime out of habit rather than analysis. Training a deep network continuously on a live stream is, per the same 2018 survey, still described as unfinished work rather than standard practice: "Despite some preliminary research, we note there are still many research challenges in this field, e.g., how to balance the tradeoff between learning accuracy, computational efficiency, learning scalability and model complexity." In practice, most deployed large language models are trained in large batches and then held fixed between releases, with any online-style updating reserved for a smaller component in front of them, such as a ranking or filtering layer.
071 min
What Online Learning Does Not Solve
Online learning does not solve the problem of dirty input. Hoi and colleagues note this directly, naming it a gap the field has largely left unaddressed: "existing online learning works seldom address the data 'veracity' issue, that is, the quality of data, which can considerably affect the efficacy of online learning." A batch pipeline usually gets a chance to clean and de-duplicate a dataset before training; an online learner updates on whatever arrives, including a mislabeled example, a bot-generated click, or a sensor glitch, the moment it shows up.
It also does not fix a learner that updates too eagerly. A model nudged too hard by every new example can overfit to a short, unrepresentative burst of similar data β a coordinated spam wave, a single viral post β and lose accuracy on the steady traffic around it once the burst passes. The usual fix is a smaller, decaying learning rate, not abandoning online updates altogether.
And it needs a real feedback signal to learn from. A system with no reliable way to observe the true outcome of each prediction β no click, no fraud confirmation, no label at all β has nothing for an online learner to update on, however fast its infrastructure runs.
081 min
Online Learning vs. Nearby Concepts
The nearest neighbor by name alone is active learning, and the two are easy to confuse because both concern data arriving over time. The deciding fact: active learning decides which unlabeled example to send for a human label next; online learning decides how to update the model once a labeled example already exists. A system can do both β actively choose what to label, then update online the moment that label comes back β but neither technique requires the other.
A second neighbor is continual learning, sometimes called lifelong learning, which is about training on a sequence of distinct tasks without forgetting earlier ones. Its concern is what the model forgets, not how often it updates. An online learner can update on every single example within one task and forget nothing, because there is only ever one task in play.
The clearest contrast is with fine-tuning: fine-tuning takes a model and continues training it on a new, but still fixed, dataset, in ordinary batch passes, then stops and ships the result. Online learning never stops taking in new examples once it is live β there is no training set that gets finished, only a stream that keeps arriving.
?6 questions
Questions people ask
What is online learning in machine learning?
Is online learning the same as active learning?
Does online learning need labeled data?
What is regret in online learning?
Can online learning replace batch training entirely?
Why don't most large language models use online learning?
Β§3 sources
Sources
Hoi, S. C. H., Sahoo, D., Lu, J., & Zhao, P. (2018). Online Learning: A Comprehensive Survey. SMU Technical Report / arXiv:1802.02871.
McMahan, H. B., Holt, G., Sculley, D., et al. (2013). Ad Click Prediction: a View from the Trenches. Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '13), 1222β1230.
Zinkevich, M. (2003). Online Convex Programming and Generalized Infinitesimal Gradient Ascent. Proceedings of the 20th International Conference on Machine Learning (ICML 2003), 928β936.





