The rapid growth of AI workloads is creating unprecedented demands on data center architectures. Modern AI applications, including large language models (LLMs), generative AI, recommendation systems, and high-performance computing, require significantly higher memory capacity, bandwidth, and efficiency.
Traditional compute-centric architectures are increasingly limited by the movement of data between processors and memory. As AI models continue to scale, excessive data movement creates performance bottlenecks, increases latency, and drives higher power consumption.
To address these challenges, the industry is moving toward memory-centric computing architectures enabled by Compute Express Link® (CXL®). CXL enables flexible memory expansion, memory pooling, and new system architectures that allow compute resources to operate more efficiently.
Marvell and SK hynix are collaborating to enable the next generation of memory-centric computing by combining Marvell® Structera™ A CXL-based near-memory acceleration technology with SK hynix advanced memory solutions. Together, Marvell and SK hynix are helping accelerate the adoption of CXL-enabled architectures for AI data centers by delivering a highly efficient and scalable approach to memory processing.
Marvell Structera™ A: Intelligent Memory Offload Engine
At the core of Structera A is a powerful CXL-based memory offload engine designed to address the growing gap between compute performance and memory capability.
Unlike traditional memory expansion solutions that primarily increase capacity, Structera A introduces intelligent processing capability closer to memory. By moving data-intensive operations closer to where data resides, Structera A reduces unnecessary data movement and improves overall system efficiency.
Structera A integrates:
This architecture enables memory-intensive processing to execute efficiently within the memory subsystem while allowing host CPUs and AI accelerators to focus on primary compute workloads.
Improving AI System Efficiency Through Memory Offload
As AI workloads continue to scale, data movement has become one of the largest contributors to system latency and power consumption. Traditional architectures require processors and accelerators to repeatedly move large datasets between compute resources and memory, reducing overall system efficiency.
Structera A addresses this challenge through intelligent workload offload by processing data closer to memory.
The Structera A architecture enables:
By moving processing closer to memory, Structera A transforms memory from a passive storage resource into an active computing element.
Marvell and SK hynix Enabling the Future of Memory-centric Computing
The combination of Marvell Structera A and SK hynix memory technology enables efficient processing for a broad range of memory-intensive applications for AI workloads.
The transition to AI-scale infrastructure requires collaboration across the entire technology ecosystem. Compute, interconnect, and memory technologies must work together to deliver scalable and efficient solutions.
Through the collaboration between Marvell and SK hynix, Structera A demonstrates how CXL-based near-memory processing can address the challenges of next-generation AI workloads. Marvell and SK hynix are enabling a new generation of AI infrastructure designed for higher performance, improved efficiency, and scalable deployment across next-generation data centers.
CXL Memory Module-Accelerator (CMM-Ax): A General-purpose CXL Processing Near Memory (CXL-PNM) to Break the Memory Wall for AI Computing
Developed in collaboration with Marvell, CMM-Ax stands as an ASIC-based CXL-PNM solution, integrating the Marvell Structera A intelligent PNM engine and SK hynix’s advanced memory and proprietary software stack. This vertically optimized architecture delivers optimal performance for diverse memory-bound workloads by enabling seamless compute offloading directly within the memory subsystem. Specifically targeting the AI domain, CMM-Ax addresses the critical bottleneck of "Long-Context" LLM inference, where conventional GPU systems struggle with memory capacity limits. By constructing an integrated inference system combining GPUs with CMM-Ax, the companies demonstrated a novel approach to efficiently accelerate these massive workloads.

The CMM-Ax Device
| CMM-Ax Specification | |
|---|---|
| Form Factor | FH3/4L AIC Full Height ¾ Length Add-in Card |
| Interconnect | CXL 2.0 / Device type 3 / PCIe 5.0 x16 |
| Memory | DDR5-6400 4-ch / 521GB / 200GB/s |
| Controller | 16 ARM Neoverse V2 cores at 3.0GHz |
Advanced PNM Platform with Vertically Co-optimized HW/SW Architecture
CMM-Ax establishes a shared coherent memory space by integrating a CXL controller, 16 Arm Neoverse v2 general-purpose cores, and four channels of DDR5 DRAM connected via CXL, eliminating data copying penalties for low-latency offloading. This hardware is orchestrated by the PNM User Library, which utilizes Memory and Command Application Programming Interfaces (APIs) to enable direct reads and writes to the CXL region, bypassing traditional interrupts and Direct Memory Access (DMA) overhead. On the device, the Device Runtime dynamically interprets commands and spawns dedicated Command Executor threads for each task, facilitating true parallel execution across the 16-core array to ensure maximum throughput and efficient resource utilization for memory-intensive workloads.

CMM-Ax HW Architecture (left) and CMM-Ax SW Stack (right). Source: SK Hynix
Overcoming Memory Bottlenecks in Long-context LLM Inference
LLM inference is inherently memory-bound, as the auto-regressive decode phase repeatedly accesses KV Cache data that grows exponentially with context length. In conventional GPU systems, surging context lengths (128K–1M tokens) exhaust on-device memory, forcing drastic batch size reductions that severely constrain throughput. CMM-Ax resolves this by leveraging CXL-based memory extension and an integrated Attention API to enable near-memory processing. By offloading overflow KV Cache to high-capacity CXL memory and executing computations directly where data resides, CMM-Ax eliminates costly data movement, allowing LLM inference systems to simultaneously support ultra-long contexts and large batch sizes with superior efficiency. To validate this approach, the companies constructed NELSSA, an integrated LLM inference system combining GPUs with CMM-Ax.
The CMM-Ax Applied LLM inference System was tested against standard GPU and CPU-offloading baselines using the Llama3-8B-1048K model1 across 128K–1024K token sequences, with accuracy validated by the RULER benchmark.2 Reported metrics reflect a projected 512 GB per-device configuration (scaled from the 32 GB prototype), focusing on decode throughput (tokens/s) as the primary efficiency metric.
The system successfully serves all workloads up to batch size 8 at 128K and 1 at 1M tokens (at a 4% selection ratio), whereas GPU-only baselines fail due to memory exhaustion. Even with prototype constraints, the Marvell and SK hynix system matches GPU baseline throughput where feasible, while projected results for the full 512 GB configuration reveal that performance gains scale with workload size as heavier loads improve PNM efficiency and amortize fixed overhead. Ultimately, the system achieves up to 5.5 times higher throughput than the single-GPU baseline and 3.6 times higher than the dual-GPU setup, proving its capability to deliver scalable, high-performance inference for ultra-long context LLMs.

CMM-Ax Applied LLM Inference System Decode Throughput. Source: SK Hynix
Together, Marvell and SK hynix are accelerating the evolution toward memory-centric computing for the AI era. SK hynix and Marvell have successfully co-developed CMM-Ax, a general-core-based ASIC CXL-PNM solution designed to support large-capacity memory, leveraging Marvell’s Structera-A CXL-PNM controller. Through this collaboration on next-generation memory solutions, SK hynix and Marvell aim to further strengthen their partnership and solidify their positions as leading global AI data center solution providers.
1. Leonid Pekelis, Michael Feil, Forrest Moret, Mark Huang, and Tiffany Peng. 2024. Llama 3 Gradient: A series of long context models. https://doi.org/10.57967/hf/3372
2. Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. 2024. RULER: What’s the real context size of your long-context language models? https://arxiv.org/pdf/2404.06654 (2024)
###
This blog contains forward-looking statements within the meaning of the federal securities laws that involve risks and uncertainties. Forward-looking statements include, without limitation, any statement that may predict, forecast, indicate or imply future events or achievements. Actual events or results may differ materially from those contemplated in this blog. Forward-looking statements are only predictions and are subject to risks, uncertainties and assumptions that are difficult to predict, including those described in the “Risk Factors” section of our Annual Reports on Form 10-K, Quarterly Reports on Form 10-Q and other documents filed by us from time to time with the SEC. Forward-looking statements speak only as of the date they are made. Readers are cautioned not to put undue reliance on forward-looking statements, and no person assumes any obligation to update or revise any such forward-looking statements, whether as a result of new information, future events or otherwise.
Tags: Company News, AI, AI infrastructure