OPSD Self-Distillation
What to do without a stronger teacher: one model plays both teacher and student – the only difference is context.
- The Teacher sees privileged information (ground truth answer, documentation) and “solves by looking at the answer” (easier than solving from scratch), outputting a word-for-word target distribution.
- The Student only sees the problem and pulls itself up to the teacher on its own trajectories.
Sources of privileged information are broader than verifiable reward.
Boundary: ceiling – the ability of the student itself; if the model cannot clearly explain the solution even with the answer in hand – there is no signal.
Related: On-Policy Distillation, RLVR and verifiable rewards