Reduce long-context attention memory
one move: fuse attention safely
Open the code pane immediately, then use the Spec tab, optional design question, tests, and artifact when you want them.
Finish each move before carrying the artifact forward.
FlashAttention tiles Q/K/V and computes softmax online without materializing the full score matrix.
Deliverable: Write the mechanism in one sentence.Implement a tiled attention kernel that fuses QKᵀ, softmax, and AV.
Deliverable: Type the smallest runnable implementation.The result must match naive attention within tolerance across sequence lengths.
Deliverable: Record the expected output and one edge case.Compare memory, tile size, and latency against the naive implementation.
Deliverable: Change one variable and explain the result.A kernel benchmark and numerical-equivalence report.
Deliverable: Save the code, output, and a short failure note.The result must match naive attention within tolerance across sequence lengths.
Compare memory, tile size, and latency against the naive implementation.
A kernel benchmark and numerical-equivalence report.
Mapped from the supplied Build Everything PDF, source page 47. The five moves are Deriva’s implementation contract for this project.