Make a language model describe images
one move: align two representations
Open the code pane immediately, then use the Spec tab, optional design question, tests, and artifact when you want them.
Finish each move before carrying the artifact forward.
A VLM projects image patch tokens into the language model's representation space.
Deliverable: Write the mechanism in one sentence.Combine a frozen vision encoder, trainable projection, and language decoder.
Deliverable: Type the smallest runnable implementation.Image-token shapes, projection dimensions, and answer grounding must be checked.
Deliverable: Record the expected output and one edge case.Change projection width, prompt format, and image resolution; inspect hallucinations.
Deliverable: Change one variable and explain the result.A vision-language model card with image-grounded examples and failures.
Deliverable: Save the code, output, and a short failure note.Image-token shapes, projection dimensions, and answer grounding must be checked.
Change projection width, prompt format, and image resolution; inspect hallucinations.
A vision-language model card with image-grounded examples and failures.
Mapped from the supplied Build Everything PDF, source page 65. The five moves are Deriva’s implementation contract for this project.