Hardware-managed heterogeneous high-bandwidth memory and flash in LLM inference systems
By Hakam Atassi , Noa Zilberman , and Amro Awad
University of Oxford

Abstract
Large Language Models (LLMs) have become the driver of many modern applications, but their parameter counts have outgrown the memory capacity of a single GPU. Because GPUs rely on High-Bandwidth Memory (HBM), whose capacity is approaching a fundamental ceiling, fitting today’s models often requires aggregating several accelerators, a cost that places edge deployment out of reach for many users. High-Bandwidth Flash (HBF) offers a denser alternative, providing 16x more capacity per stack at comparable bandwidth. In this work, we show that while replacing HBM with HBF can address the capacity problem, doing so naively severely impacts performance due to HBF’s long tail memory latency starving GPU schedulers. To address this, we propose a Heterogeneous Memory Architecture (HMA) that combines HBM and HBF through a prediction-based migration policy to keep high latency HBF off the GPU’s critical path. We demonstrate that the HMA achieves geo-mean 2.79x improved performance over an HBF Only baseline, enabling the efficient execution of 9.2x larger models on a single edge GPU.
Index Terms — High-Bandwidth Flash, Heterogeneous Memory Architecture, Large Language Model, Prefetching
To read the full article, click here
Related Semiconductor IP
- NPU IP
- JPEG XL Encoder
- I2C Master/Slave Controller Core
- NVMe Validation Test Suite
- Hybrid Memory Cube Verification IP
Related Articles
- Automated Estimation of MBIST Area and Test Time in Heterogeneous Memory IPs via Stacked Ensemble Framework
- NAND Flash memory in embedded systems
- LPDDR flash: A memory optimized for automotive systems
- The Growing Importance of AI Inference and the Implications for Memory Technology
Latest Articles
- Terracotta: Enabling the Adoption of New DRAM Techniques via a Flexible DRAM Interface and Memory Controller
- A Framework for Accelerating Transformer Inference on RISC-V for Edge AI
- An Interleaved Parallel Dependent Quantization Hardware Architecture for H.266/VVC
- A Formal Security Analysis of CAN XL
- A Secure dToF LiDAR SoC with Dual-Domain Fingerprinting and Event-Driven AFE Circuit Achieving Sensor-Level Attack Resilience