GRPO

GRPO (Group Relative Policy Optimization) dispenses with the value network: for one query, a group of trajectories is sampled, the advantage of each being its relative position within the group ((ri − mean)/std).

“Better than the group average — plus, worse — minus.”

Cheaper than PPO (no critic).

Limitation: advantage is distributed evenly across the entire answer — a coarse distribution of contribution, blurs the signal in long, multi-step tasks; there PPO with a value network or distribution at the move level.

Symmetric clipping limits the magnitude of the update per step.

Related: Reward for process and reward for result, On-Policy Distillation, Data and environment are more important than the algorithm