Train a larger model with sharded state
one move: shard what cannot fit
Open the code pane immediately, then use the Spec tab, optional design question, tests, and artifact when you want them.
Finish each move before carrying the artifact forward.
FSDP shards parameters, gradients, and optimizer state while checkpointing activations.
Deliverable: Write the mechanism in one sentence.Build a small sharded trainer with accumulation, checkpointing, and restart support.
Deliverable: Type the smallest runnable implementation.A resumed run must match an uninterrupted run within tolerance.
Deliverable: Record the expected output and one edge case.Change shard size, accumulation, and checkpoint frequency; measure memory and throughput.
Deliverable: Change one variable and explain the result.A restartable distributed-training report.
Deliverable: Save the code, output, and a short failure note.A resumed run must match an uninterrupted run within tolerance.
Change shard size, accumulation, and checkpoint frequency; measure memory and throughput.
A restartable distributed-training report.
Mapped from the supplied Build Everything PDF, source page 67. The five moves are Deriva’s implementation contract for this project.