OPSD Self-Distillation

What to do without a stronger teacher: one model plays both teacher and student – the only difference is context.

  • The Teacher sees privileged information (ground truth answer, documentation) and “solves by looking at the answer” (easier than solving from scratch), outputting a word-for-word target distribution.
  • The Student only sees the problem and pulls itself up to the teacher on its own trajectories.

Sources of privileged information are broader than verifiable reward.

Boundary: ceiling – the ability of the student itself; if the model cannot clearly explain the solution even with the answer in hand – there is no signal.

Related: On-Policy Distillation, RLVR and verifiable rewards