Write a GPU kernel that competes with torch
one move: make parallel work explicit
Open the code pane immediately, then use the Spec tab, optional design question, tests, and artifact when you want them.
Finish each move before carrying the artifact forward.
A GPU kernel maps blocks of data to thousands of parallel program instances.
Deliverable: Write the mechanism in one sentence.Implement elementwise ReLU with Triton load, compute, and store operations.
Deliverable: Type the smallest runnable implementation.Kernel output must match torch exactly across sizes and tails.
Deliverable: Record the expected output and one edge case.Change block size and compare latency against the reference implementation.
Deliverable: Change one variable and explain the result.A correctness-tested custom ReLU kernel and benchmark.
Deliverable: Save the code, output, and a short failure note.Kernel output must match torch exactly across sizes and tails.
Change block size and compare latency against the reference implementation.
A correctness-tested custom ReLU kernel and benchmark.
Mapped from the supplied Build Everything PDF, source page 32. The five moves are Deriva’s implementation contract for this project.