←
AI/ML Systems
Deriva
Learn
AI/ML
Patterns
Observe
Search
⌘K
More
⌄
0%
AI/ML Systems
0%
Home
Learn
Patterns
Observe
More
Opening…
AI/ML
›
Track
DATA-001 – BASE-005
Data engineering
0 of 15 done · 0 due for review.
all
new
started
attempted
done
all kinds
construct
predict
debug
compare
implement
operate
communicate
DATA-001
Choose the required fields for this document contract: tickets with text, label, author, and optional attachments.
construct · what-counts-as-training-data
new
DATA-002
Two rows share identical text and label. Which duplicate survives under first-occurrence-wins, and why does that rule matter?
predict · what-counts-as-training-data
new
DATA-003
A row has text: 12345 (a number). Repair this malformed record WITHOUT hiding the error.
debug · what-counts-as-training-data
new
DATA-004
Separate these three values: a missing label (absent), an unknown label ('unknown'), and an invalid label ('xqzz'). Which belongs in training?
compare · what-counts-as-training-data
new
DATA-005
Write the rejection reason for a leaked label: a row already used in the held-out set. Make the reason distinct from all other rejections.
implement · what-counts-as-training-data
new
DATA-006
Predict the result of this temporal split: tickets sorted by creation date, first 80% train, last 20% test. What risks does it carry?
predict · split-the-evidence
new
DATA-007
Author 'chen' appears in both train and test. Find the group that leaks and state the harm.
debug · split-the-evidence
new
DATA-008
Design a dataset version identifier: the same source rows cleaned twice must produce the same id, and a changed row must produce a different one.
construct · reproducibility-is-part-of-correctness
new
DATA-009
An ingestion job failed after row 400 of 1,000. Recover it from its checkpoint without reprocessing the first 400 rows.
operate · reproducibility-is-part-of-correctness
new
DATA-010
A retry resends the same ingestion batch after a timeout. Is this retry idempotent? Decide and justify.
predict · serve-the-model-safely
new
BASE-001
Build a majority-class baseline for these labels: 90 question, 6 bug, 4 feature. What does it predict for every row, and what is its accuracy?
implement · the-baseline-earns-the-right-to-improve
new
BASE-002
Build a keyword baseline for bug detection: flag rows containing 'crash', 'error', or 'broken'. State its blind spot.
construct · the-baseline-earns-the-right-to-improve
new
BASE-003
For a ranking problem (best ticket answers first), choose a baseline and predict how it performs: popularity, recency, or random.
predict · ranking-is-a-separate-responsibility
new
BASE-004
Compare this model and baseline: model 0.91, baseline 0.90, same split. Is this improvement meaningful?
compare · build-a-better-baseline-honestly
new
BASE-005
A random baseline gets 0.5 AUC. Explain to a teammate why this random baseline is still useful, not useless.
communicate · the-baseline-earns-the-right-to-improve
new