4-bit KV Caching in LMCache: Offloading Quantized KV Beyond HBM for Context-Heavy Agents on AMD MI355X

AMD ROCm integrates 4-bit KV quantization with LMCache to offload context beyond HBM limits on MI355X GPUs. The layout-aware connector preserves accuracy while moving mixed bf16 and 4-bit data to CPU DRAM. Benchmarks show improved hit rates and doubled goodput for heavy agent workloads.