← Foundry04 · Mission 1 of 4
Eight requests. One pool.
A 512 MB KV pool, eight simultaneous requests, each holding a 180–220 token conversation. Half-precision KV costs 0.5 MB per token — do the arithmetic before the queue does it for you. Quantization is a serving decision, not just a model decision.
floor controlspool 512 MB · 8 requests
KV dtype 0.5 MB/token
batch slots 8
KV allocation
batching