←
AI/ML Systems
Deriva
Learn
AI/ML
Patterns
Observe
Search
⌘K
More
⌄
0%
AI/ML Systems
0%
Home
Learn
Patterns
Observe
More
Opening…
AI/ML
›
Track
TENSOR-001 – LM-002
Deep learning and transformers
0 of 30 done · 0 due for review.
all
new
started
attempted
done
all kinds
debug
predict
implement
compare
construct
derive
operate
communicate
TENSOR-001
Repair this tensor shape mismatch: a [32, 4] batch of embeddings must multiply with weights of shape [5, 4] to produce [32, 5].
debug · tensors-are-structured-data
new
TENSOR-002
Predict the output shape of this batched operation: [32, 8] embeddings reduced over the last axis.
predict · tensors-are-structured-data
new
TENSOR-003
Implement a mean reduction over the feature axis WITHOUT losing the batch dimension: keep [32, 1] out of [32, 8].
implement · tensors-are-structured-data
new
TENSOR-004
Compare dense and sparse memory for this vocabulary: 50,000 tokens, 8 nonzero entries per row, 1 million rows.
compare · tensors-are-structured-data
new
TENSOR-005
Find the operation that causes numerical overflow: exp of logits reaching 800.
debug · tensors-are-structured-data
new
AUTOGRAD-001
Draw the computation graph for this scalar expression: L = (a·b + c)² where a, b, c are leaves.
construct · gradients-are-dependency-accounting
new
AUTOGRAD-002
Propagate one gradient through the graph: dL/dL = 1, L = u², u = a·b. What is dL/da?
derive · gradients-are-dependency-accounting
new
AUTOGRAD-003
Parameter 'b' never receives a gradient, yet it feeds the loss. Find the missing backward edge in this graph.
debug · gradients-are-dependency-accounting
new
AUTOGRAD-004
Validate this analytic gradient with finite differences: dL/dw = 2w at w = 3, h = 0.001.
implement · gradients-are-dependency-accounting
new
AUTOGRAD-005
One layer's weights never change across 100 epochs. Debug the parameter that never receives a gradient.
debug · gradients-are-dependency-accounting
new
NN-001
Choose an activation for this output contract: binary bug/question with probabilities summing to 1.
predict · compose-a-neural-network
new
NN-002
Predict how initialization changes activation scale: weights initialized 10× larger than the recommended range.
predict · compose-a-neural-network
new
NN-003
Trace one MLP forward pass: x = [1, 0], W1 = [[1, 0], [0, 1]], no bias, ReLU, output layer W2 = [1, 1].
derive · compose-a-neural-network
new
NN-004
Trace one MLP backward pass: the loss gradient at the output is 1, and the output is 1·h. What is the gradient at h?
derive · compose-a-neural-network
new
NN-005
Diagnose these exploding gradients: loss → NaN after 20 steps, weights become 1e8. Name the mechanism and two fixes.
debug · compose-a-neural-network
new
NN-006
Diagnose these vanishing gradients: early layers barely move after 200 epochs, tanh activations.
debug · compose-a-neural-network
new
NN-007
Choose where normalization changes the system: batch norm after the linear layer vs before the activation.
predict · compose-a-neural-network
new
NN-008
Resume training from this checkpoint: epoch 40 of 100, loss 0.31. What must the resume verify before continuing?
operate · compose-a-neural-network
new
NN-009
Compare batch and mini-batch updates: batch 512 vs mini-batch 16 on the same data. Predict the loss-curve difference.
compare · compose-a-neural-network
new
NN-010
Explain why this model memorizes training data: 40M parameters, 30K training rows, accuracy 1.0/0.62 train/test.
communicate · compose-a-neural-network
new
ATTN-001
Construct queries, keys, and values for this sentence: 'the model failed'. One token, one vector each.
construct · context-changes-representation
new
ATTN-002
Calculate one attention score: Q1 = [1, 0], K2 = [0, 1], scale = √2. What is the pre-softmax score between token 1 and token 2?
derive · context-changes-representation
new
ATTN-003
Apply this causal mask to a 3-token attention matrix: tokens may only attend to themselves and earlier tokens.
implement · context-changes-representation
new
ATTN-004
Predict how context changes this token representation: 'bank' appears first with 'river' and then with 'money'.
predict · context-changes-representation
new
ATTN-005
Inspect this attention failure: the model attends to the padding token at the end of every sentence. Explain the mechanism and the harm.
communicate · context-changes-representation
new
ATTN-006
Compare one head and multiple heads: single-head attention captures one relationship per position. What do multiple heads add?
compare · context-changes-representation
new
ATTN-007
Explain the cost of a longer context: attention is all-pairs. How does the compute scale with context length n?
derive · context-changes-representation
new
ATTN-008
Debug this positional-information bug: the model treats 'dog bites man' and 'man bites dog' as identical.
debug · context-changes-representation
new
LM-001
Calculate next-token cross-entropy: the true next token is 'bug', the model assigns it probability 0.1.
derive · build-a-tiny-language-model
new
LM-002
Compare temperature and top-k sampling: temperature 1.5 vs 0.2, and top-k 5 vs 50. Predict the style of each generation.
compare · build-a-tiny-language-model
new