Foundry
Six stations, six real engines. You do not read about modern AI systems — you operate them: train the tokenizer, turn the sampling dials, read the attention arcs, run the serving floor, fix the retrieval bench, and guard the agent loop. Every station is the actual machinery, small enough to hold.
Token Bench
Train a real byte-pair encoding tokenizer on a corpus, watch merges appear one by one, then budget and price prompts by the token.
resume · mission 1 →Sampling Deck
A language model with its probabilities exposed. Temperature, top-k and nucleus truncation reshape the distribution in front of you — learn what each dial really does before you ever set one in production.
Attention Lens
One attention head with the covers off: scaled dot products, softmax weights, positional mixing. Predict where attention lands, then fix a context bug the way you fix prompts.
Serving Floor
An inference server as a discrete-event simulation. Requests arrive, KV cache fills, batches form. Choose dtype, batching and paging — then keep the SLA when traffic doubles.
Retrieval Bench
A RAG pipeline that fails the way real ones fail: vocabulary mismatch, keyword stuffing, chunk boundaries cutting answers in half. Diagnose, then repair with chunking and query expansion.
Agent Loop
A ReAct agent you can actually break: runaway retries, observation bloat blowing the context, prompt injection hiding in a file. Set the guardrails, run the loop, read the trace.
Production AI work is full of invisible machinery — tokenizers, samplers, attention, KV caches, retrievers, agent loops. You can memorize their names, or you can operate them. Every station here is the real algorithm at bench scale: deterministic, inspectable, and small enough to break on purpose.