SFT First, Then RL

RL requires parsable output: if the model outputs gibberish instead of JSON, the reward function can’t score anything – there’s nothing to learn from.

SFT acts as “first learn to speak clearly”: a small number of demonstrations stabilizes the format. The reverse order doesn’t work – the reward signal turns into noise.

Rule boundaries: with a sufficiently strong base model, you can go straight to RL (DeepSeek-R1-Zero), but at the cost of readability – that’s why R1 still added a cold start SFT.

Form first, then spirit.

Related: Three stages of post-training, SFT memorizes - RL generalizes