Fine-Tuning

Continuing to train a model that already learned general patterns during pretraining on a smaller, task-specific dataset, so it adapts to a narrower job without starting from scratch.

12 min read

Β· Also in

By Ravi SuranaUpdated 3 sources

Quick answer

~20 sec

Fine-tuning is taking a model already trained on a large, general dataset and continuing its training on a smaller, task-specific dataset, so it adapts to a narrower job without starting from scratch. It exists because training a large model from nothing is too expensive for most teams. It differs from prompting, which changes the input, not the model.

011 min

Fine-Tuning at a Glance

  • What it is β€” continuing to train a pretrained model on a smaller, task-specific dataset.
  • Why it exists β€” training a large model from scratch is out of reach for almost every team.
  • What it costs β€” full fine-tuning updates every parameter, so the result is as large as the original model.
  • When it breaks β€” a fine-tuned model can overfit a small dataset and forget general capability it had before.
  • As of 2026 β€” parameter-efficient methods like LoRA are the default over full fine-tuning for most teams, not a niche shortcut.

021 min

The Problem Fine-Tuning Solves

Before fine-tuning was standard practice, the naive approach to a new NLP task was building and training a model architecture specific to that task from randomly initialized weights, using only whatever labeled data existed for that exact problem. That approach fails whenever labeled data is scarce, which is most of the time: a task like classifying support tickets by urgency might have a few thousand labeled examples, nowhere near enough to teach a model language itself from nothing, only enough to teach it the specific task on top of language it already understands.

Jacob Devlin and colleagues at Google framed the shift plainly in the paper that introduced BERT: the pretrained model can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks, such as question answering and language inference, without substantial task-specific architecture modifications. The expensive, data-hungry part, learning general language patterns, happens once, on a huge general-purpose dataset; the cheap, fast part, adapting to one specific job, happens per task, on whatever labeled data that task actually has.

032 min

How Fine-Tuning Works

Picture a support-ticket triage tool built on a general-purpose language model. Before fine-tuning, the base model already understands English, code, and a huge range of general knowledge, from pretraining on a broad corpus of text. It has never specifically learned this company's product names, its specific urgency categories, or the tone its support team uses. Fine-tuning takes that general model and continues training it, using ordinary gradient-descent updates, on a dataset of this company's own tickets paired with the correct urgency label, so the model's weights shift toward this specific task without losing the general language ability learned during pretraining.

The traditional version of this process, full fine-tuning, updates every one of the model's parameters during that second training pass, which means the result is a complete second copy of the model, as large as the original, for every task fine-tuned. Edward Hu and colleagues at Microsoft identified the resulting deployment problem directly: the major downside of fine-tuning is that the new model contains as many parameters as in the original model, which becomes a critical deployment challenge once a model reaches the scale of GPT-3's 175 billion parameters, since deploying independent instances of fine-tuned models, each with 175B parameters, is prohibitively expensive.

Their proposed fix, called LoRA, freezes the pretrained model's original weights entirely and instead trains a small set of added matrices, injected into each layer, that capture just the change needed for the new task. Because those added matrices are dramatically smaller than the full model, LoRA can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times compared to full fine-tuning, while performing on par with full fine-tuning on standard benchmarks.

041 min

A Concrete Example

Fine-tuning is not only used to teach a new task; it is also the main technique used to change how a model behaves toward its users, not just what it knows. Long Ouyang and colleagues at OpenAI fine-tuned GPT-3 in two stages, first on human-written demonstrations of the desired behavior, then further on human rankings of model outputs using reinforcement learning, to produce a model family they called InstructGPT. The scale of the effect was large enough to be counterintuitive: in human evaluations on their prompt distribution, outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. A model roughly a hundred times smaller, fine-tuned specifically to follow instructions the way people actually wanted, beat a much larger model that had only been pretrained.

052 min

What This Means for Your Work

For a founder or PM deciding how to build an AI feature, fine-tuning changes the build-versus-prompt decision: if the task can be described clearly enough in a prompt with a handful of examples, prompting a general model is faster to ship and cheaper to iterate on, and fine-tuning is worth the added cost mainly when the task needs a consistent tone, a specialized vocabulary, or a behavior that a prompt alone keeps failing to reliably produce across many inputs.

For an engineer, fine-tuning changes what data pipeline needs to exist before a model can ship: a labeled, task-specific dataset, a held-out validation slice to catch overfitting, and a decision between full fine-tuning and a parameter-efficient method like LoRA, which changes both the training cost and how many separate task-specific models can realistically be hosted at once.

For a designer or content lead, fine-tuning is the lever that actually changes a model's voice: a model fine-tuned on a specific brand's writing samples will default toward that tone without needing the tone re-specified in every single prompt, the way a prompted-only approach requires.

Across all three roles, the decision fine-tuning changes is really the same one: whether the team is willing to own a training pipeline and a versioned model artifact, or would rather keep behavior entirely in prompts and retrieved context, which is easier to change but harder to make perfectly consistent.

061 min

What Fine-Tuning Costs

The clearest cost of full fine-tuning is storage and deployment, not just training compute: a separately fine-tuned copy of a large model is exactly as large as the original, so a team running ten fine-tuned variants of the same base model is paying to store and potentially serve ten full copies. Hu and colleagues' benchmark on GPT-3 175B makes the gap concrete: full fine-tuning with Adam against LoRA's parameter-efficient approach differs by roughly 10,000 times in trainable parameters and 3 times in GPU memory required for training, which is why LoRA and similar methods became the default rather than a niche optimization once model sizes crossed into the tens of billions of parameters.

There is also a data cost that does not show up in a compute bill: fine-tuning needs labeled, task-specific examples, and InstructGPT's own approach required OpenAI to collect both human-written demonstrations and human rankings of outputs, a genuinely expensive, ongoing labeling effort rather than a one-time dataset purchase.

Teams weighing this tradeoff also give up some flexibility: a fine-tuned model is versioned and frozen at a point in time, while a prompt can be edited and redeployed in minutes with no retraining at all.

071 min

What Fine-Tuning Does Not Solve

Fine-tuning does not solve every adaptation problem, and using it where a lighter technique would do is a common, expensive mistake. A task that changes daily, such as answering questions about today's inventory or this week's pricing, is the wrong target for fine-tuning entirely, since a fine-tuned model's knowledge is frozen at the point training stopped; that kind of constantly-changing information belongs in a retrieval step at answer time, not baked into model weights that would need retraining on every update.

Fine-tuning can also fail on its own terms. A model fine-tuned on too little or too narrow a dataset can overfit it, memorizing the specific examples rather than the general pattern behind the task, the same overfitting failure that affects any trained model. A separate, well-documented risk is catastrophic forgetting: aggressive fine-tuning on a narrow task can measurably degrade a model's general capabilities that existed before fine-tuning began, trading broad competence for narrow competence more than the task actually required.

A fine-tuned model can also drift out of date the same way any trained model does: a support-ticket classifier fine-tuned on last year's product lineup needs a fresh fine-tuning pass, or a switch to retrieval, once the product lineup changes, since the fine-tuning run itself has no way to notice the world moved on.

081 min

Fine-Tuning vs. Nearby Concepts

The nearest neighbor is transfer learning, and the relationship is one of category and instance rather than two competing techniques: transfer learning is the general strategy of reusing knowledge a model learned on one task or dataset for a different, related task, and fine-tuning is the specific technique of continuing to train that model's weights to do it. Every fine-tuning run is an act of transfer learning; not every form of transfer learning involves fine-tuning; a model can also transfer what it learned through prompting alone, with no weight updates at all.

A second neighbor is prompt engineering, and the deciding fact between them is whether the model's weights change. Prompt engineering changes only the input given to a fixed, unmodified model, so it is fast to iterate and produces no separate model artifact to store or deploy. Fine-tuning changes the model itself, which costs more upfront but can produce more consistent behavior across inputs than a prompt alone reliably achieves, especially at the volume a production feature actually sees.

091 min

A Second Case

BERT and LoRA show fine-tuning solving two different problems a decade apart. BERT's authors used fine-tuning to adapt one pretrained model architecture to eleven different NLP tasks with only a small added output layer per task, obtaining new state-of-the-art results across all eleven, including pushing the GLUE benchmark score to 80.5%, a 7.7 point absolute improvement over the prior best result, and SQuAD v1.1 question answering to a Test F1 of 93.2. The problem BERT's fine-tuning solved was breadth: one general architecture, adapted cheaply to many specific tasks.

LoRA's authors, working with a model roughly five hundred times larger than BERT, faced a different problem entirely: not whether fine-tuning worked, but whether it was affordable to deploy at all once full fine-tuning meant storing a complete 175-billion-parameter copy for every task. Their fix did not change what fine-tuning accomplishes; it changed what fine-tuning costs, which is why the technique became necessary once model scale, not task variety, became the binding constraint.

101 min

Where the Evidence Is Contested

Not every researcher treats fine-tuning as the right default lever for adapting model behavior. Ouyang and colleagues' own results carried a caveat alongside the headline finding: they reported improvements in truthfulness and reductions in toxic output generation while having minimal performance regressions on public NLP datasets, the word "minimal" doing real work β€” fine-tuning for one goal, following instructions well, was not entirely free with respect to the model's prior general capability, even in the paper credited with showing the technique's biggest win.

A second, ongoing debate concerns whether parameter-efficient methods like LoRA are a full substitute for full fine-tuning or a close approximation with its own gaps. Hu and colleagues reported LoRA performs on-par or better than fine-tuning in model quality on the benchmarks they tested, but on-par on a benchmark is not the same claim as identical in every deployment; teams choosing between the two are trading a well-quantified efficiency gain against a less precisely quantified risk that a narrower, added-matrix approach captures the target behavior slightly less completely than updating every parameter would.

?6 questions

Questions people ask

Is fine-tuning the same as training a model from scratch?

No. Fine-tuning starts from a model that already learned general patterns during pretraining and continues training it on a smaller, task-specific dataset, which needs far less data and compute than training from randomly initialized weights.

When should I use fine-tuning instead of a well-crafted prompt?

Use fine-tuning when a task needs a consistent tone, specialized vocabulary, or a behavior a prompt keeps failing to produce reliably across many inputs. If a prompt with a few examples already works, prompting is faster to ship and cheaper to iterate on.

What are the risks or limits of fine-tuning?

A model can overfit a small fine-tuning dataset, or suffer catastrophic forgetting, losing general capability it had before fine-tuning began. It also can't fix knowledge that changes daily, since a fine-tuned model's knowledge is frozen at training time.

Does fine-tuning require GPUs and a lot of compute?

Full fine-tuning updates every parameter and needs real compute and storage, scaling with model size. Parameter-efficient methods like LoRA reduce trainable parameters by roughly 10,000 times and GPU memory by 3 times on large models, at similar quality.

How much does fine-tuning cost compared to full fine-tuning versus LoRA?

Full fine-tuning produces a complete second copy of the model per task, as large as the original. LoRA trains a much smaller set of added parameters instead, cutting both training cost and storage substantially while performing on par on standard benchmarks.

Is fine-tuning the same as transfer learning?

Fine-tuning is one specific technique within the broader strategy of transfer learning. Every fine-tuning run transfers knowledge from pretraining to a new task, but transfer learning can also happen without any weight updates, through prompting alone.

Β§3 sources

Sources

  1. Devlin, J., Chang, M-W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805.

  2. Hu, E. J., Shen, Y., Wallis, P., et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685.

  3. Ouyang, L., Wu, J., Jiang, X., et al. (2022). Training Language Models to Follow Instructions with Human Feedback. arXiv:2203.02155.

Keep reading

More from AI

All of AI
All of AI