By Chander Chadha, Director of Product Marketing, Storage Products, ³Ô¹ÏºÚÁÏ
As AI models grow larger and inference workloads scale, larger models, longer context windows and growing KV caches are driving demands on memory resources. Traditional CPU-centric networked JBOF (Just Bunch of Flash) can¡¯t keep pace with the throughput, latency, and efficiency requirements of these modern AI clusters.
DPU (data processing unit) -based storage is a compelling alternative to traditional storage, acting as the broker between the network and SSD for remote storage. By offloading storage, networking and security processing from the host CPU onto a dedicated DPU, AI infrastructure can move data closer to compute, reduce latency, and free up valuable CPU cycles for AI workloads.
To address this need, ³Ô¹ÏºÚÁÏ offers the OCTEON DPU family, purpose-built for hyperscale cloud workloads and data center applications, extending its use specifically into network storage acceleration for AI environments.
By Khurram Malik, Associate Vice President, Custom Cloud Solutions, ³Ô¹ÏºÚÁÏ
![]()
The rapid growth of AI is creating unprecedented demand for memory capacity. As models get larger and workloads become more data-intensive, traditional server architectures are increasingly challenged to scale memory efficiently.
At Flash Memory Summit (FMS), ³Ô¹ÏºÚÁÏ demonstrated how CXL can enable a more flexible and scalable approach to memory infrastructure, showcasing an end-to-end architecture at the ? booth.
The demonstration combined an Intel platform with three ³Ô¹ÏºÚÁÏ? technologies: Structera? X CXL memory expander, Structera? S CXL switch and Alaska? P PCIe retimer. Together, they demonstrated how memory expansion, CXL switching and high-speed PCIe connectivity can work together within a real server platform.
By Arifur Rahman, Director of Product Marketing, Custom Cloud Solutions, ³Ô¹ÏºÚÁÏ
Memory has become the defining constraint of modern AI infrastructure. Large language models, in-memory databases, and deep learning recommendation models all share the same bottleneck: there is never enough DRAM. Compute Express Link (CXL) was designed to break that bottleneck but a CXL memory expander is only as valuable as the breadth of platforms it can run on.
That is why ecosystem enablement is a core pillar of ³Ô¹ÏºÚÁÏ's CXL strategy. Today we are marking a new milestone: successful interoperability of the ³Ô¹ÏºÚÁÏ? Structera? X CXL memory-expansion controller with the .
By Khurram Malik, AVP, Data Center Memory and Storage Solutions, ³Ô¹ÏºÚÁÏ
CXL has become one of the most important technologies shaping AI infrastructure. As hyperscalers race to deploy larger AI models, longer context windows and increasingly memory-intensive inference workloads, memory capacity and bandwidth have emerged as critical constraints on performance, efficiency and scaling. At the same time, CXL adoption is reaching an inflection point, moving from evaluation into real-world deployment across hyperscale environments.
³Ô¹ÏºÚÁÏ is leading this transition with Structera? X memory expansion solutions developed alongside the world¡¯s leading hyperscalers. The story begins with the shipping of Structera X 2404 and 2504 platforms, which have enabled hyperscalers to expand memory resources more efficiently, including extending the useful life of existing DDR4 investments while powering demanding AI workloads.
Structera X is not a series of disconnected product eras¡ªit is a single, continuous architectural evolution. Today¡¯s generation is already delivering real hyperscaler deployments, ecosystem maturity and a compelling TCO advantage. From that foundation, ³Ô¹ÏºÚÁÏ is extending the architecture toward the next phase of AI infrastructure innovation, adding capabilities enabled by the evolving CXL 3.2 ecosystem, PCIe Gen 6 connectivity and more advanced multi-host memory sharing architectures. These advancements will create larger, more flexible memory pools, enabling more efficient sharing of resources across servers and improving infrastructure utilization at hyperscale. As the architecture advances, ³Ô¹ÏºÚÁÏ is driving it toward higher bandwidth, deeper data optimization and increasingly disaggregated memory environments built to meet the growing demands of AI workloads.
By Ravi Mahatme, Senior Director, Product Management, Photonic Fabric Business Unit, ³Ô¹ÏºÚÁÏ
AI infrastructure is entering a new architectural era. As AI inference scales, performance increasingly depends not simply on adding more compute, but on how efficiently compute, memory and connectivity operate together as one unified AI infrastructure system. Modern AI inference workloads are insatiable consumers of memory. Large language models, reasoning models and agentic AI applications require rapid access to model parameters, embeddings and rapidly growing key-value (KV) caches that preserve conversational context. These working data sets are growing into the hundreds of gigabytes, and increasingly terabytes, making memory capacity, bandwidth and latency just as important as accelerator performance. As AI infrastructure scales, overall system performance increasingly depends on how efficiently accelerators can access, move and utilize memory resources rather than simply adding more compute.
Why AI Needs a New Memory Tier
Today's AI memory hierarchy was never designed for inference at the scale modern workloads demand. High bandwidth memory (HBM) attached directly to GPUs delivers exceptional performance but remains expensive and capacity constrained. System DRAM provides larger memory pools but cannot economically scale alongside every accelerator. NVMe SSDs offer abundant capacity, yet their latency makes them unsuitable for serving active inference workloads.
This challenge is especially visible in large language models, where growing KV caches must remain readily accessible to avoid repeatedly recomputing previous tokens. Keeping these caches entirely in HBM is prohibitively expensive, while moving them to storage introduces latency that reduces token generation performance. The result is that GPUs increasingly spend valuable cycles waiting for data rather than performing inference.