Implement the verifiable core of grouped policy optimization: subtract the group mean reward so updates express relative advantage.
Required APIdef relative_advantages(rewards):
return centered_rewardsBehaviorThink through the mechanism first if you want the extra reasoning step. It never blocks the editor.
relative_advantages([1.0, 3.0])the better sample receives positive relative advantage
relative_advantages([2.0, 2.0])no sample wins when evidence is tied
2 hidden edge tests run after the visible contract passes.
Run Tests to see the contract verdicts here.