KL - mode-seeking vs mass-covering

KL divergence is asymmetric, and direction matters.

  • Reverse KL KL(πθ‖πref) penalizes the policy moving to where the reference model gives almost zero – it is mode-seeking: concentrates probability on a few best peaks and discards the rest.
  • Forward KL is mass-covering: forces covering all patterns of the reference.

This explains why responses after RL sound more confident and less diverse: the model keeps high-quality strategies and decisively discards the rest.

Maximum likelihood SFT is pure mass-covering.

Related: RLHF, SFT memorizes - RL generalizes