← Projectsdata-pipeline / L4Rebuild from Manifest~20 min
one move: bound a failure
Make raw text batchable by defining a deterministic lowercase whitespace tokenizer before adding a learned vocabulary. In the Data Pipeline system, implement this level as a deterministic contract before composing it with the next service.
Required API
def tokenize(text):
return token_list
Behavior
normalize case
split on runs of whitespace
return tokens in input order
Examples — tap to reveal
Constraints
Use only the Python standard library.
Return deterministic values so every run can be compared with the contract.