Make a chatbot prefer better answers
one move: optimize for preferences
Open the code pane immediately, then use the Spec tab, optional design question, tests, and artifact when you want them.
Finish each move before carrying the artifact forward.
SFT teaches instruction following; DPO shifts preference toward chosen answers without a reward model.
Deliverable: Write the mechanism in one sentence.Fine-tune adapters on demonstrations, then optimize chosen/rejected pairs with DPO.
Deliverable: Type the smallest runnable implementation.Preference loss, held-out comparisons, and refusal behavior must be visible.
Deliverable: Record the expected output and one edge case.Vary preference margins, beta, and data quality; inspect regressions.
Deliverable: Change one variable and explain the result.An aligned checkpoint, preference report, and safety evaluation.
Deliverable: Save the code, output, and a short failure note.Preference loss, held-out comparisons, and refusal behavior must be visible.
Vary preference margins, beta, and data quality; inspect regressions.
An aligned checkpoint, preference report, and safety evaluation.
Mapped from the supplied Build Everything PDF, source page 69. The five moves are Deriva’s implementation contract for this project.