RLVR and Verifiable Rewards
RLVR – Reinforcement Learning with Verifiable Rewards: the signal comes from a rule-checker (did the test pass, is the answer correct) rather than a trained model.
Agent tasks are mostly verifiable – thus RLVR is the main line. A special case of the general scheme: the reward model can be deterministic code when correctness is definable by a rule.
Asymmetry of verification and generation (“it’s easier to check than to do”) – the fundamental reason RL can outperform any demonstrator.
Related: Reward for process vs reward for outcome, RLHF, Data and environment are more important than the algorithm