SFT Memorizes - RL Generalizes
- SFT optimizes “looks like the reference” (maximum likelihood) – learns a rigid input→output mapping and memorizes: applies the old answer when the environment rules change.
- RL optimizes “how good the outcome is” (expected reward) – learns a transferable strategy and generalizes: applies the same process to new conditions.
GeneralPoints experiment: when the rules change (J/Q/K = 11/12/13 instead of 10) SFT scaling drops by 5–8%, RL scaling grows by +3–17%.
Root: mass-covering (SFT covers all patterns) vs mode-seeking (RL concentrates on the best peaks).
Related: Three stages of post-training, SFT then RL, On-Policy Distillation