011 min
RLHF at a glance
- What it is: Trains a model to produce the responses human raters prefer, not just likely text.
- Why it exists: Predicting the next word does not teach a model to follow instructions or refuse a bad request.
- What it costs: Thousands of paid human comparisons, plus a separate reward model to train.
- When it breaks: A model can learn to sound convincing rather than be correct β reward hacking.
- As of: As of 2026, RLHF and Direct Preference Optimization are both used across production chat models.
021 min
The problem RLHF solves
A language model trained only to predict the next word in a large body of text learns to continue writing in whatever style and content that text used. The training objective does not distinguish a genuinely helpful, honest answer from a plausible-sounding one that is evasive, made up, or beside the point β both continue the text equally well by that measure. Collecting example answers and applying ordinary fine-tuning narrows this gap: the model now imitates specific good responses. But a fixed set of examples cannot cover every question a user will ask, and it gives the model no signal about the many other ways a response can go wrong that nobody wrote a correct answer for. What is missing is a way to state, in general, that one kind of answer is better than another, across responses nobody specifically demonstrated, and to measure "better" by what real people actually judge as an improvement rather than by what a single demonstration writer happened to write down. Reinforcement Learning from Human Feedback provides that: it trains a separate model to predict which of two responses a person would prefer, then uses that prediction to reward better responses during further training β turning alignment with human intent into something the model can be trained against directly, instead of a fixed list of examples.
032 min
How RLHF works
RLHF trains a model in three separate stages, each producing something the next stage uses. Take one running example: a base language model is being adapted to summarize a customer's complaint email into three sentences a support agent can scan quickly.
1. Supervised fine-tuning
The base model β already trained to predict the next word over a large body of text β is fine-tuned on a smaller set of examples: real complaint emails paired with a three-sentence summary a human writer produced. This step, called supervised fine-tuning (SFT), teaches the model the general shape of the task β write a short summary, not a poem or a refusal β but it only ever sees one correct answer per email. It has no way to learn that one imperfect summary is closer to correct than another.
2. Reward model
To give the model a sense of degrees, not just right and wrong, the SFT model generates several different summaries of the same email. A person compares two of those summaries at a time and picks the better one β not by writing a fresh summary, but by choosing between ones the model already produced. Thousands of these paired comparisons train a second, separate model β the reward model β to output a single number predicting which of two summaries a human would prefer. The reward model never writes text; its only job is to score it.
3. Reinforcement learning against the reward model
The original SFT model β now called the policy β generates a new summary, the reward model scores it, and a reinforcement learning algorithm adjusts the policy's weights to make higher-scoring summaries more likely. Christiano and colleagues, and later Ouyang and colleagues, both used an algorithm called Proximal Policy Optimization for this step. A penalty term keeps the policy's outputs close to what the original SFT model would have written, so the policy cannot drift into text that scores well on the reward model but reads as strange or repetitive to an actual person. This third stage is the one RLHF is named for, and the one where the real improvement over supervised fine-tuning alone happens: the model is now optimizing directly for what raters said they preferred, across far more variation than any fixed set of examples could cover.
041 min
A concrete example
Ouyang and colleagues (2022) trained InstructGPT by applying RLHF's three stages to GPT-3, then had human labelers compare its answers with plain GPT-3's answers to the same prompts, without being told which model produced which response. Ouyang and colleagues found that outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. The size of the model mattered less than whether it had been trained to answer the way the prompt actually asked.
The same pattern shows up in something as simple as one instruction.
Before RLHF (base model): asked to explain compound interest to a beginner, a model trained only to predict a likely next word often continues into a plausible-sounding tangent β a history of interest rates, a comparison of account types, a disclaimer that it is not financial advice β none of which the instruction asked for.
After RLHF (instruction-tuned model): the same request gets what the instruction actually asked for: a definition, one worked number, and a stop. The model does not know more about compound interest than it did before. Reward-model training rated "answers what was asked, briefly" higher than "is a plausible continuation," and the reinforcement-learning stage made the higher-scoring response more likely.
051 min
How RLHF changed since
Christiano, Leike, Brown, Martic, Legg, and Amodei (2017) first used RLHF's three-stage structure to train reinforcement-learning agents to play Atari games and control simulated robots, using comparisons between short video clips of the agent's behavior instead of a hand-written reward function. Christiano and colleagues needed feedback on fewer than one percent of the agent's interactions with its environment to train it.
Stiennon and colleagues (2020) moved the same structure to text, training a reward model on human comparisons between article summaries and using it to fine-tune a summarization policy. On Reddit's TL;DR dataset, Stiennon and colleagues found that their model's summaries significantly outperformed both human reference summaries and much larger models fine-tuned with supervised learning alone.
Ouyang and colleagues (2022) then applied the same recipe to instruction-following in general, not just summarization, producing InstructGPT and fixing supervised fine-tuning, reward modeling, and reinforcement learning as the standard three-stage pipeline behind the current generation of chat models.
Rafailov and colleagues (2023) then asked whether the reinforcement-learning stage was necessary at all. Rafailov and colleagues introduced Direct Preference Optimization, which solves the standard RLHF problem with only a simple classification loss instead of a separate reward model and reinforcement-learning step.
062 min
What RLHF means for your work
For an engineer choosing how to align a model to a product's own definition of a good answer, RLHF is the reason teams collect human-in-the-loop feedback instead of only writing more prompt instructions: a well-run RLHF pass changes what the model does across an entire range of prompts, where a system prompt only changes what it does when that exact instruction is present. That is also the concrete decision it forces: before starting an RLHF pass, decide what "better" means precisely enough that two different human raters would pick the same response most of the time. A reward model trained on inconsistent preferences learns an inconsistent notion of quality, and every later stage inherits that inconsistency.
For a founder or PM deciding whether to run RLHF at all, the real cost is not the training compute β it is the ongoing work of collecting and re-collecting human comparisons every time the product's definition of a good answer changes. That is why most teams outside the largest AI labs use a vendor's already-tuned model and add a system prompt or a small fine-tuning pass on top, rather than running the full three-stage pipeline themselves. Running an in-house reward model only pays off once the product has a specific, high-volume notion of quality that a general-purpose model's existing tuning does not already cover β a support tone a general model gets wrong in a specific, correctable way, for example.
For a designer or engineer evaluating whether a change to a model's behavior actually worked, the same logic applies to testing: comparing two candidate responses side by side and picking the better one is a more reliable way to judge a change than asking a single rater to score one response in isolation, because it is the same kind of comparison judgment the reward model itself was trained on.
071 min
What RLHF costs
Collecting human comparisons is the largest recurring cost, because it does not stop after one round: every change to what counts as a "better" answer means collecting fresh comparisons, and the labelers doing that comparing are paid for their time. Training the reward model and running reinforcement learning on top of it adds a full second and third training run beyond ordinary fine-tuning, plus the compute to generate several candidate responses per prompt for the reward model to rank during training.
None of this adds cost to answering a single user request afterward. The reward model and the reinforcement-learning stage are used only during training, so a deployed RLHF-tuned model responds at the same speed as one trained only with supervised fine-tuning. The ongoing cost sits entirely in maintaining the training pipeline, not in serving it.
082 min
What RLHF does not solve
RLHF optimizes a model to produce the response human raters scored highest, and rated preference is not the same thing as truth. A rater comparing two responses is judging which one sounds more correct, more complete, or more confident in the time they have to read it β and a fluent, confident, wrong answer regularly beats a correctly hedged one on that basis. This is a known limitation of the method, not a corner case in one model: training can teach a model to sound more certain than its training data supports, because certainty was what got rewarded.
RLHF does not fix hallucination for the same reason. A fabricated fact stated fluently and a true fact stated fluently look identical to a reward model that was never shown which one was actually correct β only which one a rater preferred reading. The common overclaim is that RLHF makes a model "aligned" in some general sense. What it actually does is make the model's outputs closer to what the specific raters in the specific training rounds preferred, which is a narrower and more fragile result than "aligned" suggests.
Reward hacking is the sharpest version of the same failure: given enough reinforcement-learning steps against a fixed reward model, a policy can find responses that score well on that reward model specifically without being good responses in any wider sense. That is one reason RLHF training is stopped early rather than run to convergence.
092 min
RLHF vs. nearby concepts
| Concept | What it does | How it differs from RLHF |
|---|---|---|
| Supervised fine-tuning | Trains a model to imitate a fixed set of human-written examples | Sees only one correct answer per example; RLHF adds a preference signal for responses nobody wrote a correct answer to |
| Direct Preference Optimization (DPO) | Trains directly on the same human comparison data RLHF uses | Skips the separate reward model and the reinforcement-learning stage, solving the same preference-alignment problem with one classification loss instead of three stages |
| Reinforcement learning (general) | Learns behavior that maximizes a reward signal, usually a programmed one such as a game score | RLHF's reward comes from a model trained on human preferences, not from a rule anyone wrote down |
RLHF is most often confused with plain supervised fine-tuning, because both start with the same fine-tuning step and both need labeled data. The difference is what the label captures: a fine-tuning example is one correct answer someone wrote out in full; a preference comparison is a judgment between two answers the model already produced, which is cheaper to collect at scale and covers far more of the ways an answer can go wrong.
RLHF is also confused with Direct Preference Optimization, since DPO was built to reach the same result β a policy aligned with rated preferences β while removing the separate reward model and the reinforcement-learning stage described above. The deciding fact is whether a separate reward model and a distinct reinforcement-learning stage exist: RLHF has three stages, DPO folds them into one.
102 min
Frequently asked questions about RLHF
Is RLHF the same as fine-tuning?
No. Supervised fine-tuning trains a model on a fixed set of correct examples. RLHF adds a further stage on top of that: a reward model trained on human comparisons, then reinforcement learning that pushes the model toward whichever responses that reward model scores highest.
What is the difference between RLHF and DPO?
RLHF trains a separate reward model and then runs reinforcement learning against it, in two distinct stages. Direct Preference Optimization (Rafailov et al., 2023) skips both: it trains the policy directly on the same human comparison data using a single classification loss.
Does RLHF require a GPU cluster to run?
Training does β it needs compute for three separate stages: fine-tuning, reward-model training, and reinforcement learning. Answering a single user request afterward costs no more than a fine-tuned model, since the reward model runs only during training, not at inference.
Why does RLHF need a separate reward model?
Because a reinforcement-learning algorithm needs a score to optimize, and no one writes a numeric score for every possible response by hand. The reward model is trained once, on human comparisons between responses, so it can supply that score automatically for any new response the policy generates during training.
Can RLHF make a model sound more confident than it should?
Yes. RLHF rewards whichever response human raters preferred, and raters often prefer confident-sounding answers over correctly hedged ones. A model can learn that confident phrasing scores well regardless of whether the claim is true β a known limitation of the method, not a bug in one model.
How much human feedback does RLHF need?
Enough paired comparisons to train a reliable reward model β typically thousands, collected from multiple raters comparing pairs of responses to the same prompt. The exact number depends on how consistent the raters are and how varied the task is.
Is RLHF still used now that DPO exists?
Yes. DPO is a lighter-weight alternative for the same preference-alignment goal, not a replacement that made RLHF obsolete. Ouyang and colleagues' three-stage recipe is still the pipeline several major chat models were trained with, and both approaches are in active use.
?7 questions
Questions people ask
Is RLHF the same as fine-tuning?
What is the difference between RLHF and DPO?
Does RLHF require a GPU cluster to run?
Why does RLHF need a separate reward model?
Can RLHF make a model sound more confident than it should?
How much human feedback does RLHF need?
Is RLHF still used now that DPO exists?
Β§4 sources
Sources
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., & Amodei, D. (2017). Deep Reinforcement Learning from Human Preferences. arXiv:1706.03741.
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., & Christiano, P. (2020). Learning to Summarize from Human Feedback. arXiv:2009.01325.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training Language Models to Follow Instructions with Human Feedback. arXiv:2203.02155.
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). Direct Preference Optimization: Your Language Model Is Secretly a Reward Model. arXiv:2305.18290.





