Reward for Process and Reward for Outcome

Key choice for multi-step tasks:

  • reward for outcome – only at the end: maximum freedom to explore, but sparse signal;
  • reward for process – each step: easier learning, but risk of limiting innovation.

Outcome-and-Process Reward (RLVP) compromise: R = O + β·Φ, where Φ is a verifiable path signal: penalty for each verifiable violation and partial reward for achievable progress.

Dense signal rescues groups of solid failures (zero variance in GRPO → no gradient).

Penalty – universally applicable; reward for progress – only when progress is achievable.

Related: Four Principles of Rubric, RLVR and Verifiable Rewards, GRPO