←
AI/ML Systems
Deriva
Learn
AI/ML
Patterns
Observe
Search
⌘K
More
⌄
0%
AI/ML Systems
0%
Home
Learn
Patterns
Observe
More
Opening…
AI/ML
›
Track
MATH-001 – THEORY-006
Mathematics and theory
0 of 30 done · 0 due for review.
all
new
started
attempted
done
all kinds
predict
derive
debug
compare
counterexample
construct
communicate
transfer
MATH-001
A batch of 32 support tickets each has 4 features. Three new features are added to every row. What shape does the batch have now?
predict · tensors-are-structured-data
new
MATH-002
Why does a dot product rank document A higher than document B for this query, when they share the same vocabulary size?
derive · tensors-are-structured-data
new
MATH-003
Compute the cosine similarity of these two vectors. Which of the three options is correct?
predict · similarity-is-a-representation-choice
new
MATH-004
This matrix product is rejected at runtime. Which multiplication is invalid, and what is the rule that catches it?
debug · tensors-are-structured-data
new
MATH-005
A vocabulary of 20,000 tokens is stored as one-hot rows. A sparse representation uses only the 8 nonzero positions per row. Which is smaller, and by roughly how much?
predict · tensors-are-structured-data
new
MATH-006
This broadcast fails. Repair it without changing the batch dimension: add a text column to [32, 4] using a [32] vector of lengths.
debug · tensors-are-structured-data
new
STAT-001
A survey is sent only to users who visited the settings page last week. Which sampling process does this create, and what does it bias?
predict · the-baseline-earns-the-right-to-improve
new
STAT-002
90% of tickets are 'question', 10% are 'bug'. A model predicts 'question' for every ticket. What is its accuracy — and why is that number dangerous?
predict · the-baseline-earns-the-right-to-improve
new
STAT-003
From this confusion matrix, calculate precision, recall, and F1 for the bug class.
derive · when-a-good-score-lies
new
STAT-004
Two confidence intervals estimate the same mean: [0.42, 0.58] and [0.47, 0.53]. Which is narrower, and what caused the difference?
compare · when-a-good-score-lies
new
STAT-005
Users who open the onboarding email also finish onboarding at higher rates. Construct a reason why this correlation is NOT evidence that the email causes completion.
counterexample · the-baseline-earns-the-right-to-improve
new
STAT-006
A model says 0.9 for every prediction, yet it is wrong on half of them. Diagnose this calibration failure and name the fix.
debug · when-a-good-score-lies
new
OPT-001
Estimate the derivative of f(x) = x² at x = 3 using a finite difference with h = 0.01.
derive · learn-by-reducing-error
new
OPT-002
Derive the gradient of this one-parameter squared loss L = (w·x − y)² with respect to w.
derive · learn-by-reducing-error
new
OPT-003
Trace one gradient-descent update by hand: w = 1.0, learning rate 0.1, gradient 2.0. What is the next w?
predict · learn-by-reducing-error
new
OPT-004
A run converges at learning rate 0.01. Predict the effect of doubling the learning rate to 0.02 on this same convex loss.
predict · the-learning-rate-is-a-system-choice
new
OPT-005
In this loss trace, the loss explodes to NaN after step 40. Find the exploding update and name the cause.
debug · learn-by-reducing-error
new
OPT-006
A batch must fit in 512 MB. Each row costs 4 KB of memory per feature and there are 10 features. Choose the largest safe batch size and justify it.
construct · the-learning-rate-is-a-system-choice
new
INFO-001
Calculate the entropy of these two label distributions: A = {0.5, 0.5} and B = {0.9, 0.1}.
derive · build-a-tiny-language-model
new
INFO-002
The true label is class 0; the model outputs 0.99 for class 1. Why does cross-entropy punish this confident wrong prediction so hard?
derive · build-a-tiny-language-model
new
INFO-003
Tokenize 'run-run' under the vocabulary ['run', 'un', '-', 'run-'] with longest-match-first. Show your token sequence.
construct · build-a-tiny-language-model
new
INFO-004
A tokenizer is given a 12-character string but returns 14 tokens. Find the tokenization that creates this inflated context and fix it.
debug · build-a-tiny-language-model
new
INFO-005
A dataset has 20,000 rare words across 5,000 rows. Choose sparse bag-of-words or dense embeddings and justify your choice for a small classifier.
compare · features-are-a-decision
new
INFO-006
A dimensionality reduction of ticket texts to 2 dimensions lost the bug/question distinction. Identify what the reduction discarded and write the lesson for the team.
communicate · similarity-is-a-representation-choice
new
THEORY-001
A feature 'is_from_training_batch' is added to every row to help the model. Find the train/test leakage in this feature and state the harm.
counterexample · what-counts-as-training-data
new
THEORY-002
Which of these two models will overfit the training set, and why: a 2-parameter linear model or a 200-parameter polynomial on 30 rows?
predict · regularization-is-a-tradeoff
new
THEORY-003
Construct a distribution shift that breaks this model: trained on weekday ticket traffic, served on weekends.
construct · find-the-error-family
new
THEORY-004
Validation loss rises while training loss keeps falling. Choose the intervention: more data, L2 regularization, or more epochs — and justify.
predict · regularization-is-a-tradeoff
new
THEORY-005
A team adds 10× more data but the bug class still underperforms. Explain why more data may not fix this bias, and write the finding.
communicate · build-a-better-baseline-honestly
new
THEORY-006
Design a split that respects both time and user groups: sessions arrive hourly, and the same user recurs daily.
transfer · split-the-evidence
new