Drop in a bottleneck — “Llama 70B at 4.2s/token on A100”. WARP spins up 4 hostile optimizers that fight over quantization, kernels, and KV-cache until you get a shippable plan.
“SDXL 12s/image on A10G”. We parse model, hardware, batch, SLO.
Simulates roofline, memory vs compute bound, KV pressure, kernel breakdown.
Quant vs Kernel vs Systems vs Cost argue. Best trade wins, not loudest.
Diff, Dockerfile, Triton kernels, and cost math. Copy-paste to prod.
Don't quantize everything. Quantize the bottleneck layers, fuse the rest, and let speculation cover tail latency.