RLHF

RLHF – Reinforcement Learning from Human Feedback.

Three-stage InstructGPT pipeline:

  1. SFT on demonstrations;
  2. Reward Model on human pairwise comparisons (Bradley-Terry: comparing “which is better” is more reliable than absolute ratings);
  3. PPO with RM score.

KL penalty keeps policy close to reference model: don’t stray too far, or RM score is unjustified extrapolation. Inverse KL – mode-seeking: preserves multiple peaks of high reward.

Reward Model overoptimization: surrogate reward grows, real quality initially grows, then falls (Goodhart’s Law).

Related: KL - mode-seeking versus mass-covering, Process reward vs outcome reward, Three stages of post-training