Supermicro, Vast Data, and Solidigm Outline New Storage Tier for Agentic AI Inference
Expanding context windows and key-value cache demands are pushing tech vendors to create intermediate storage layers to reduce reliance on costly GPU memory.

As enterprise deployments transition from initial model training toward autonomous agentic workloads, underlying storage infrastructure is becoming a central planning consideration for infrastructure architects. Agentic systems perform multi-step reasoning, execution, and continuous reassessment, a process that continuously generates data and expands the active context window required during inference.
Executive insights shared during the Supermicro Open Storage Summit, first reported by SiliconANGLE, highlighted how these expanding key-value (KV) caches necessitate specialized data management tiers. Panelists from Solidigm Inc., Vast Data Inc., and Super Micro Computer Inc. detailed how combining high-performance solid-state drives, rack-scale integration, and AI-focused data software can relieve operational pressure on graphics processing units.
Scott Shadley, director of technology planning at Solidigm, noted that managing application latency during inference requires evaluating both volumetric data growth and the physical speed at which information moves across system components. He explained that modern Non-Volatile Memory Express (NVMe) SSDs have facilitated the creation of a middle tier—often described as a 3.5 layer—to house context data that previously could not be stored outside main memory.
To support this hierarchy, Solidigm provides components such as its high-speed D7-PS1010 SSD alongside high-capacity options like the D5-P5336 drive. Supermicro integrates these drives into hardware systems designed to catch memory spillovers as KV cache flows out from high-bandwidth memory (HBM) and system RAM toward local and networked storage.
Ben Lee, director of solution management at Supermicro, outlined the company's Context Memory eXtension (CMX) initiative, which formalizes this G3.5 storage tier. Lee emphasized that resolving hardware bottlenecks for large clusters demands a rack-level approach, as scaling up high-bandwidth memory on GPUs remains cost-prohibitive for standard enterprise budgets.
From the software perspective, Vast Data Director of AI Architecture Anat Heilper emphasized that caching key-value data enables organizations to trade storage resources for expensive compute capability. High KV cache hit rates allow workloads to bypass redundant processing, which simultaneously trims hardware costs and shrinks response latency.
Heilper referenced practical testing using Nvidia Dynamo, where the offloading architecture delivered a 20-fold acceleration in time-to-first-token metric for end users alongside a 90 percent decrease in required GPU compute time. These metrics underscore the performance gains achieved when context memory is efficiently managed across dedicated storage layers.
Panel participants noted that storage architectures will continue to evolve as enterprise requirements shift, meaning single static designs will not fit every inference deployment. As a result, closer technical alignment between hardware component providers, system integrators, and software vendors will be required to build custom AI infrastructure moving forward.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.

