←
AI/ML Systems
Deriva
Learn
AI/ML
Patterns
Observe
Search
⌘K
More
⌄
0%
AI/ML Systems
0%
Home
Learn
Patterns
Observe
More
Opening…
AI/ML
›
Track
EXP-001 – EVAL-010
Experiments and evaluation
0 of 20 done · 0 due for review.
all
new
started
attempted
done
all kinds
communicate
construct
debug
implement
predict
derive
EXP-001
Write a one-sentence hypothesis before changing this feature: replace the bigram feature with a trigram feature.
communicate · the-learning-rate-is-a-system-choice
new
EXP-002
Select the one variable allowed to change in this experiment: learning rate, feature set, seed, and split are all candidates.
construct · the-learning-rate-is-a-system-choice
new
EXP-003
Baseline scored 0.88 on split A, model scored 0.90 on split B. Detect this invalid comparison and state the rule it broke.
debug · build-a-better-baseline-honestly
new
EXP-004
Reproduce this run from its config and seed: seed 42, lr 0.01, 100 epochs. What must the runner fix to make replay exact?
implement · reproducibility-is-part-of-correctness
new
EXP-005
Read this learning curve: train and val loss both fall slowly and plateau high. Diagnose underfitting and predict the fix.
predict · regularization-is-a-tradeoff
new
EXP-006
Read this learning curve: train loss near zero, val loss rising after epoch 20. Diagnose overfitting and predict the fix.
predict · regularization-is-a-tradeoff
new
EXP-007
An ablation gains +0.002 F1 at 3× serving cost. Decide whether this improvement is practically meaningful and justify.
predict · build-a-better-baseline-honestly
new
EXP-008
Design an ablation for this feature group: text length, word count, and punctuation ratio are added together.
construct · build-a-better-baseline-honestly
new
EXP-009
Classify these errors by root cause: too-short tickets, rare vocabulary, duplicate training rows, and a leaked eval row.
implement · find-the-error-family
new
EXP-010
From this error taxonomy — 60% vocabulary errors, 25% duplicates, 15% short-text — choose the next experiment and predict its effect.
predict · find-the-error-family
new
EVAL-001
Choose metrics for this high-recall moderation system: missing a bad comment costs 50× a false flag. Which metric must be maximized?
predict · the-baseline-earns-the-right-to-improve
new
EVAL-002
Calculate PR-AUC from these points: recall (0.2, precision 1.0), (0.5, 0.8), (0.9, 0.5).
derive · the-baseline-earns-the-right-to-improve
new
EVAL-003
A model reports 0.95 accuracy but the 'feature' class recall is 0.1. Find the metric that hides this minority-class failure.
debug · when-a-good-score-lies
new
EVAL-004
Design slice evaluation for these user segments: mobile, desktop, and API-only users.
construct · find-the-error-family
new
EVAL-005
Interpret this calibration curve: predictions in the 0.9 bin are correct 60% of the time. Write the finding.
communicate · when-a-good-score-lies
new
EVAL-006
Build a regression fixture for this previous failure: a bug was missed because it contained no bug keywords.
implement · reproducibility-is-part-of-correctness
new
EVAL-007
Decide whether this model is ready for shadow traffic: offline F1 0.82, regression fixtures pass, calibration ECE 0.04.
predict · measure-the-service-not-just-the-model
new
EVAL-008
Latency rose 40% while quality held. Separate data, model, and serving regressions: which bucket does this belong to?
debug · measure-the-service-not-just-the-model
new
EVAL-009
Explain why offline success did not transfer online: offline F1 0.9, live precision 0.6. Write the likely causes.
communicate · measure-the-service-not-just-the-model
new
EVAL-010
Write the acceptance question for this model release: bug recall must beat the baseline on the fixed split, with no slice below 0.5.
communicate · the-baseline-earns-the-right-to-improve
new