Summarize an inference workload with warmup exclusion, p95 latency, throughput, quality, and SLO verdicts.
Required APIdef benchmark_report(runs, quality_target, p95_limit_ms, throughput_target):
return {'runs': ..., 'p95_ms': ..., 'throughput': ..., 'quality_ok': ..., 'slo_ok': ...}BehaviorThink through the mechanism first if you want the extra reasoning step. It never blocks the editor.
benchmark_report([{'latency_ms': 10, 'tokens': 10, 'quality': 1.0, 'warmup': True}, {'latency_ms': 20, 'tokens': 20, 'quality': 0.9, 'warmup': False}, {'latency_ms': 30, 'tokens': 30, 'quality': 0.9, 'warmup': False}], 0.8, 40, 900)quality and performance must pass together
benchmark_report([{'latency_ms': 10, 'tokens': 10, 'quality': 0.5, 'warmup': False}], 0.8, 20, 1)['slo_ok']a faster but worse model is not a benchmark win
2 hidden edge tests run after the visible contract passes.
Run Tests to see the contract verdicts here.