Hardware-managed heterogeneous high-bandwidth memory and flash in LLM inference systems
By Hakam Atassi , Noa Zilberman , and Amro Awad
University of Oxford

Abstract
Large Language Models (LLMs) have become the driver of many modern applications, but their parameter counts have outgrown the memory capacity of a single GPU. Because GPUs rely on High-Bandwidth Memory (HBM), whose capacity is approaching a fundamental ceiling, fitting today’s models often requires aggregating several accelerators, a cost that places edge deployment out of reach for many users. High-Bandwidth Flash (HBF) offers a denser alternative, providing 16x more capacity per stack at comparable bandwidth. In this work, we show that while replacing HBM with HBF can address the capacity problem, doing so naively severely impacts performance due to HBF’s long tail memory latency starving GPU schedulers. To address this, we propose a Heterogeneous Memory Architecture (HMA) that combines HBM and HBF through a prediction-based migration policy to keep high latency HBF off the GPU’s critical path. We demonstrate that the HMA achieves geo-mean 2.79x improved performance over an HBF Only baseline, enabling the efficient execution of 9.2x larger models on a single edge GPU.
Index Terms — High-Bandwidth Flash, Heterogeneous Memory Architecture, Large Language Model, Prefetching
To read the full article, click here
Related Semiconductor IP
- TSMC 7nm 0V75 / 0V9 ESD Local Clamp – Low Cap
- TSMC 65nm 3V3 ESD Local Clamp – Rad Hard
- TSMC 5nm 1V8, 1.2V and 0.9V ESD Local Protection – Low Cap
- TSMC 3nm 3V3 ESD Local Clamp
- TSMC 3nm 1V2 ESD Local Clamp – Low Capacitance
Related Articles
- Automated Estimation of MBIST Area and Test Time in Heterogeneous Memory IPs via Stacked Ensemble Framework
- NAND Flash memory in embedded systems
- LPDDR flash: A memory optimized for automotive systems
- The Growing Importance of AI Inference and the Implications for Memory Technology
Latest Articles
- Hardware-managed heterogeneous high-bandwidth memory and flash in LLM inference systems
- LACE: Large Language Model Aided Multi-Agent Framework for Agile RISC-V Instruction Extension
- A Process-Aware Hybrid Si/IGO Monolithic-3D 6T SRAM with BEOL Pass-Gates for the 2nm Node
- Automated Estimation of MBIST Area and Test Time in Heterogeneous Memory IPs via Stacked Ensemble Framework
- VIPER: Architecture-Aware Performance Modeling for Processing-in-Memory Design-Space Exploration