Train a model that checks its own answers
one move: learn from verifiable rewards
Open the code pane immediately, then use the Spec tab, optional design question, tests, and artifact when you want them.
Finish each move before carrying the artifact forward.
GRPO samples groups of completions, verifies answers, and updates relative advantage without a value model.
Deliverable: Write the mechanism in one sentence.Build a small verifiable-reward loop with grouped sampling and policy updates.
Deliverable: Type the smallest runnable implementation.Rewards must be reproducible, answers verifiable, and policy updates bounded.
Deliverable: Record the expected output and one edge case.Change group size, reward shaping, and verifier strictness; inspect reward hacking.
Deliverable: Change one variable and explain the result.A reasoning-training report with verifier and reward traces.
Deliverable: Save the code, output, and a short failure note.Rewards must be reproducible, answers verifiable, and policy updates bounded.
Change group size, reward shaping, and verifier strictness; inspect reward hacking.
A reasoning-training report with verifier and reward traces.
Mapped from the supplied Build Everything PDF, source page 71. The five moves are Deriva’s implementation contract for this project.