Turn a 224×224 image into 196 patch tokens
one move: turn pixels into tokens
Open the code pane immediately, then use the Spec tab, optional design question, tests, and artifact when you want them.
Finish each move before carrying the artifact forward.
A Vision Transformer treats fixed-size image patches as a sequence of tokens.
Deliverable: Write the mechanism in one sentence.Normalize an image, split it into patches, flatten, and project each patch.
Deliverable: Type the smallest runnable implementation.Patch count, dimension, and ordering must match the image geometry.
Deliverable: Record the expected output and one edge case.Change patch size and inspect how sequence length and detail change.
Deliverable: Change one variable and explain the result.A patch manifest and visual tokenization report.
Deliverable: Save the code, output, and a short failure note.Patch count, dimension, and ordering must match the image geometry.
Change patch size and inspect how sequence length and detail change.
A patch manifest and visual tokenization report.
Mapped from the supplied Build Everything PDF, source page 43. The five moves are Deriva’s implementation contract for this project.