Understanding Its Purpose First
RLHF uses human feedback on outcomes to shape model behavior; the common pipeline converts preferences into rewards and then trains the model.
Understanding Through an Example
For the same lesson-planning question, humans find Answer A more grade-appropriate while Answer B is too vague. Gathering many such comparisons helps establish evaluation signals, then the model is trained to favor responses that receive higher rewards.
Which ones better meet the requirements
Typical pipeline for training reward model
For example, using PPO
Going Deeper
InstructGPT is a classic example: first use demonstrations for supervised fine-tuning, then train a reward model with human rankings, followed by reinforcement learning. Here, human feedback is typically first converted into automatically computable rewards, eliminating the need for humans to score each generated token in real time.
What It Cannot Explain
RLHF is a class of training pipelines, and PPO is a method that can be used within it. DPO uses preference pairs for direct optimization, offering another path to understanding preference training. Rewards may drift from their intended target, and higher rewards do not directly prove factual correctness.
Where to Go Next
References
- Training language models to follow instructions with human feedback: InstructGPT example: demonstration data, human rankings, post-training with feedback.
- Proximal Policy Optimization Algorithms: Sections 2–3: policy gradient, probability ratios and clipped surrogate objectives; clipping does not impose hard constraints on all probability changes.
- Direct Preference Optimization: Direct objective on preference pairs; standard form does not require a separate reward model or PPO online sampling loop.
References verified on 2026-09-09; original papers are used to illustrate mechanisms, and examples in the text are for instructional purposes.