RLHF
RLHF – Reinforcement Learning from Human Feedback.
Three-stage InstructGPT pipeline:
- SFT on demonstrations;
- Reward Model on human pairwise comparisons (Bradley-Terry: comparing “which is better” is more reliable than absolute ratings);
- PPO with RM score.
KL penalty keeps policy close to reference model: don’t stray too far, or RM score is unjustified extrapolation. Inverse KL – mode-seeking: preserves multiple peaks of high reward.
Reward Model overoptimization: surrogate reward grows, real quality initially grows, then falls (Goodhart’s Law).
Related: KL - mode-seeking versus mass-covering, Process reward vs outcome reward, Three stages of post-training