Treat GPU memory like virtual memory
one move: allocate blocks, not illusions
Open the code pane immediately, then use the Spec tab, optional design question, tests, and artifact when you want them.
Finish each move before carrying the artifact forward.
Paged attention manages variable-length KV state with fixed blocks, free lists, and explicit fragmentation instead of one giant contiguous allocation.
Deliverable: State the system invariant in one sentence.Allocate and free fixed-size blocks for requests, report logical-to-physical mappings, and reject capacity overflow.
Deliverable: Type the deterministic core.Freed blocks return to the pool, mappings preserve token order, and two requests never alias a live block.
Deliverable: Record a normal case and a boundary case.Free a middle request, allocate a longer request, and compare fragmentation against contiguous allocation.
Deliverable: Name the failure signal and its guard.A KV page table, allocator trace, and memory-utilization report.
Deliverable: Save the contract, evidence, and handoff note.Freed blocks return to the pool, mappings preserve token order, and two requests never alias a live block.
Free a middle request, allocate a longer request, and compare fragmentation against contiguous allocation.
A KV page table, allocator trace, and memory-utilization report.
Deriva-authored extension beyond the supplied PDF. The five moves are Deriva’s implementation contract for this project.