On-Policy Distillation

On-Policy Distillation combines the strengths of SFT and RL:

  • the student generates trajectories themselves (on-policy — solves the distribution mismatch problem);
  • a stronger teacher provides a word-by-word distribution at each step (dense signal instead of sparse reward at the end).

~10x fewer training steps than pure RL.

The teacher evaluates not “right/wrong”, but the complete distribution “with what probability each token”.

Requirement: a sufficiently realistic environment for the student to explore in the deployment distribution.

Related: OPSD Self-Distillation, GRPO, Three Stages of Post-Training