Express the memory-saving idea behind FlashAttention as a tiled causal reduction: compute each output from a bounded trailing window without materializing a score matrix.
Required APIdef flash_prefix_average(values, tile):
return outputsBehaviorThink through the mechanism first if you want the extra reasoning step. It never blocks the editor.
flash_prefix_average([1.0, 3.0, 5.0], 2)a bounded tile still preserves local causal context
flash_prefix_average([2.0, 4.0], 1)a one-token tile is the identity
2 hidden edge tests run after the visible contract passes.
Run Tests to see the contract verdicts here.