Shorter technical notes on inference serving, GPU optimization, and platform engineering.
Why autoregressive serving has two distinct performance regimes — and why conflating them leads to wrong capacity plans.
How context length, batch size, and GQA interact to determine whether your cluster is weight-bound or cache-bound.
A practical comparison for teams choosing a production inference stack — latency, throughput, and operational trade-offs.
Multi-step tool use changes the request distribution. What that means for GPU sizing, caching, and tail latency.