Increasing Tokens-per-Dollar with CXL: FMS 2026 Recap

3 min read

The CXL Consortium returned to Future of Memory and Storage (FMS) 2026, August 4–6 in Santa Clara, California, to showcase how CXL technology can help AI infrastructure increase tokens-per-dollar by expanding available memory, reducing storage bottlenecks and enabling more efficient use of memory resources.

From video demos at the CXL Consortium kiosk and an expert panel on AI inference to new CXL 3.x solutions and demonstrations from members across the show floor, FMS highlighted the growing momentum behind CXL technology.

CXL Consortium Kiosk at the Open Standards Pavilion

At the CXL Consortium kiosk in the Open Standards Pavilion, attendees connected with Consortium representatives and explored how CXL enables organizations to increase tokens-per-dollar through more efficient memory architectures, improved resource utilization and scalable AI infrastructure.

The kiosk featured video demonstrations from CXL Consortium members, highlighting CXL solutions in action across AI inference, memory expansion, persistent memory and disaggregated memory. The video demonstrations showed how the CXL ecosystem is putting expanded, pooled and disaggregated memory into practice. By increasing available memory and keeping more workload data closer to compute, CXL can reduce dependence on slower storage tiers, improve resource utilization and help organizations get more from their AI infrastructure.

View the FMS 2026 CXL Video Demos

CXL Innovation Recognized at FMS

CXL ecosystem innovation was also recognized through the FMS 2026 Best of Show Awards, with CXL 4.0 and CXL Memory Pooling receiving the Connectivity Award.

The recognition reflects the continued innovation across the CXL ecosystem as companies develop the connectivity, memory and infrastructure technologies needed for increasingly data-intensive workloads.

CXL Experts Explore How to Increase AI Infrastructure Efficiency

The CXL Consortium-sponsored panel, “Beyond the Memory Wall: How CXL Memory Pooling and Sharing Are Transforming AI Inference,” brought together experts in memory architecture and AI infrastructure to examine how CXL can address some of the most pressing challenges facing AI inference deployments.

Moderated by Anil Godbole (Intel), CXL Consortium Marketing Working Group Chair, the panel featured Sandeep Dattaprasad (Astera Labs), Sumit Puri (Liqid), Luis Ancajas (Micron) and Ronen Hyatt (UnifabriX).

As AI models and context windows continue to grow, so does the amount of KV Cache data that needs to be stored and accessed during inference. The panel explored how CXL memory pooling and sharing can provide additional memory capacity closer to compute, along with memory tiering and KV Cache offloading strategies that reduce reliance on slower storage.

For KV Cache workloads, reducing the storage bottleneck can accelerate access to the data needed for inference while improving the utilization of valuable compute resources. CXL memory pooling also creates opportunities for CPUs, GPUs, and accelerators to access and share memory resources more efficiently across heterogeneous infrastructure.

CXL Ecosystem Momentum on Display

The momentum behind CXL extended throughout the FMS show floor, where Consortium members showcased technologies based on CXL 3.x and demonstrated how the ecosystem continues to mature across memory devices, controllers, pooling architectures and software.

Several demonstrations highlighted large-scale memory expansion, including solutions capable of adding terabytes of DDR5 capacity through CXL. These examples showed how data centers can increase available memory independently of CPU-attached DRAM, creating greater flexibility to support AI, analytics, in-memory databases and other memory-intensive applications.

Other demonstrations focused on CXL 3.x memory pooling and disaggregation, including multi-host memory configurations and dynamic memory allocation. These capabilities show how CXL can help infrastructure operators provision memory more efficiently and make capacity available where it is needed, rather than statically attaching it to individual processors.

AI inference was another major theme. Demonstrations from across the ecosystem showed how CXL pooled and expanded memory can support KV Cache storage, transfer and reuse, creating larger memory tiers for increasingly demanding AI workloads. Approaches combining different memory and storage technologies also demonstrated how CXL can enable tiered architectures that balance capacity, performance and cost.

The ecosystem also showcased continued progress in CXL 3.2 controllers and 64 GT/s connectivity, validation tools, persistent memory, processing near memory and composable infrastructure.

Collectively, these developments show that CXL is progressing across the full ecosystem, from silicon and controllers to memory devices, software and complete system architectures.

Looking Ahead for 2026

As AI inference and emerging agentic AI applications require larger memory footprints, moving data between limited local memory and slower storage can constrain valuable compute resources. CXL provides new options to expand, pool and share memory across CPUs, GPUs and accelerators, creating more flexible memory architectures for these demanding workloads.

By expanding available memory, reducing storage bottlenecks, improving memory utilization and enabling new approaches to KV Cache tiering and offload, CXL can help organizations get more value from their compute infrastructure and increase tokens-per-dollar.

Thank you to all the members, speakers, and attendees who joined us at FMS 2026 and helped demonstrate the CXL ecosystem’s continued momentum. We look forward to seeing many of you at the AI Infra Summit event in Santa Clara, from September 15 – 17.

Facebook
Twitter
LinkedIn