Tiered KV Cache Offloading in vLLM
vLLM introduces tiered KV cache offloading to preserve evicted data across host memory, storage, and remote peers. This feature reduces recomputation, lowers latency, and increases effective serving capacity for long-context models. The framework enables horizontal cache scaling and warm-starting new instances from shared storage, available since version 0.22.