Hardware-managed heterogeneous high-bandwidth memory and flash in LLM inference systems
By Hakam Atassi , Noa Zilberman , and Amro Awad
University of Oxford

Abstract
Large Language Models (LLMs) have become the driver of many modern applications, but their parameter counts have outgrown the memory capacity of a single GPU. Because GPUs rely on High-Bandwidth Memory (HBM), whose capacity is approaching a fundamental ceiling, fitting today’s models often requires aggregating several accelerators, a cost that places edge deployment out of reach for many users. High-Bandwidth Flash (HBF) offers a denser alternative, providing 16x more capacity per stack at comparable bandwidth. In this work, we show that while replacing HBM with HBF can address the capacity problem, doing so naively severely impacts performance due to HBF’s long tail memory latency starving GPU schedulers. To address this, we propose a Heterogeneous Memory Architecture (HMA) that combines HBM and HBF through a prediction-based migration policy to keep high latency HBF off the GPU’s critical path. We demonstrate that the HMA achieves geo-mean 2.79x improved performance over an HBF Only baseline, enabling the efficient execution of 9.2x larger models on a single edge GPU.
Index Terms — High-Bandwidth Flash, Heterogeneous Memory Architecture, Large Language Model, Prefetching
To read the full article, click here
Related Semiconductor IP
- Mesochronous Bridge for PCIe/CXL
- AI-native GPU
- OpenGMSL Verification IP
- OpenGMSL Leaf IP
- FlexGen Multi-Die Smart Network-on-Chip (NoC) IP
Related Articles
- Automated Estimation of MBIST Area and Test Time in Heterogeneous Memory IPs via Stacked Ensemble Framework
- NAND Flash memory in embedded systems
- LPDDR flash: A memory optimized for automotive systems
- The Growing Importance of AI Inference and the Implications for Memory Technology
Latest Articles
- FlexSpIM: An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Hybrid Stationarity
- MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators
- Analog Pin Directionality as an Exfiltration Attack Surface in Mixed-Signal ICs
- SIMT-Aware Lockstep Verification and Functional-Coverage Closure Methodology for an Open-Source RISC-V GPGPU: A UVM 1.2 Environment
- Efficient Hardware Information-Flow Tracking for Pre-Silicon Security Testing