Small enough to replay every accept/reject decision in the browser worker.
Duplicates and the leak must be findable, testable, and reproducible across runs.
What counts as training data?
Today’s one move: predict which rows survive
Someone hands you a CSV of support tickets. 30 rows. 'Train a model to route them,' they say. Before you write a single line of model code, something else has to happen first — and most pipelines skip it.
Commit before anything is explained: which row is the most dangerous to train on?
The rows are lying to you
Two rows are identical. One row has no label at all. One row isn't even a row — the text field is missing, and another one's 'text' is a number. And one row was already used in last week's evaluation export. A model doesn't care about any of this. It will happily learn from every one of them — and then quietly fail on real tickets.
The real question
What would you require a row to be, before you allowed it near a model? That requirement is a contract. This lesson is about writing it — before any training, before any evaluation, before any service.