RLVR and Verifiable Rewards

RLVR – Reinforcement Learning with Verifiable Rewards: the signal comes from a rule-checker (did the test pass, is the answer correct) rather than a trained model.

Agent tasks are mostly verifiable – thus RLVR is the main line. A special case of the general scheme: the reward model can be deterministic code when correctness is definable by a rule.

Asymmetry of verification and generation (“it’s easier to check than to do”) – the fundamental reason RL can outperform any demonstrator.

Related: Reward for process vs reward for outcome, RLHF, Data and environment are more important than the algorithm