Know when a confident prediction is wrong
one move: measure the right failure
Open the code pane immediately, then use the Spec tab, optional design question, tests, and artifact when you want them.
Finish each move before carrying the artifact forward.
A single average score hides class imbalance and confidence errors; evaluation needs confusion counts and calibrated probabilities.
Deliverable: State the invariant in one sentence.Compute accuracy, precision, recall, F1, and expected calibration error from deterministic predictions and confidence buckets.
Deliverable: Type the smallest deterministic implementation.Metrics must handle empty classes, preserve input order, and expose the confusion matrix rather than only one headline number.
Deliverable: Record a normal case and a boundary case.Swap the positive class, add an overconfident wrong prediction, and compare the metric deltas.
Deliverable: Name the failure signal and the guard you would add.An evaluation card with metric definitions, confusion counts, and a calibration note.
Deliverable: Save the contract, evidence, and handoff note.Metrics must handle empty classes, preserve input order, and expose the confusion matrix rather than only one headline number.
Swap the positive class, add an overconfident wrong prediction, and compare the metric deltas.
An evaluation card with metric definitions, confusion counts, and a calibration note.
Deriva-authored extension beyond the supplied PDF. The five moves are Deriva’s implementation contract for this project.