As Large Language Model (LLM) inference scales across multi-GPU deployments, the Key-Value (KV) Cache becomes a major performance bottleneck. Each GPU typically keeps its own isolated cache, which can spill to host DRAM. When different GPUs process the same prompt prefix, each one repeats the same prefill computation. CXL fabric-attached memory (FAM) addresses this by providing multiple GPUs with a shared, byte-addressable memory pool, allowing other GPUs to reuse KV Cache entries computed by one GPU.
During this webinar, CXL Consortium member Micron will share lessons from integrating CXL FAM-based KV Cache sharing into an open-source multi-GPU inference framework. The session will focus on the solution architecture and the practical value it delivers:
- A block-level cache manager that allocates in CXL memory
- Sequence-hash-based cache discovery across GPUs
- An asynchronous transfer engine that moves blocks between CXL and GPU memory without stalling the inference scheduler
- The CUDA host-registration workflow that enables direct memory access (DMA) between GPUs and CXL memory
Early results show that CXL FAM performs on par with a host DRAM baseline. Attendees will leave with practical insight into how CXL memory pooling and sharing can improve AI inference efficiency at scale.
Speaker: Skyler Windh, MTS Systems Software Engineering, Micron Technology
Register for the webinar HERE.