Let different heads learn different patterns
one move: specialize heads
Open the code pane immediately, then use the Spec tab, optional design question, tests, and artifact when you want them.
Finish each move before carrying the artifact forward.
Several attention heads run in parallel over separate representation subspaces.
Deliverable: Write the mechanism in one sentence.Split dimensions into heads, apply attention, concatenate, and project.
Deliverable: Type the smallest runnable implementation.Head shapes must divide evenly and concatenation must restore model width.
Deliverable: Record the expected output and one edge case.Vary the number of heads and compare per-head attention patterns.
Deliverable: Change one variable and explain the result.A multi-head attention module with head-level diagnostics.
Deliverable: Save the code, output, and a short failure note.Head shapes must divide evenly and concatenation must restore model width.
Vary the number of heads and compare per-head attention patterns.
A multi-head attention module with head-level diagnostics.
Mapped from the supplied Build Everything PDF, source page 27. The five moves are Deriva’s implementation contract for this project.