Hardware-managed heterogeneous high-bandwidth memory and flash in LLM inference systems

By Hakam Atassi , Noa Zilberman , and Amro Awad 
University of Oxford

Abstract

Large Language Models (LLMs) have become the driver of many modern applications, but their parameter counts have outgrown the memory capacity of a single GPU. Because GPUs rely on High-Bandwidth Memory (HBM), whose capacity is approaching a fundamental ceiling, fitting today’s models often requires aggregating several accelerators, a cost that places edge deployment out of reach for many users. High-Bandwidth Flash (HBF) offers a denser alternative, providing 16x more capacity per stack at comparable bandwidth. In this work, we show that while replacing HBM with HBF can address the capacity problem, doing so naively severely impacts performance due to HBF’s long tail memory latency starving GPU schedulers. To address this, we propose a Heterogeneous Memory Architecture (HMA) that combines HBM and HBF through a prediction-based migration policy to keep high latency HBF off the GPU’s critical path. We demonstrate that the HMA achieves geo-mean 2.79x improved performance over an HBF Only baseline, enabling the efficient execution of 9.2x larger models on a single edge GPU.

Index Terms — High-Bandwidth Flash, Heterogeneous Memory Architecture, Large Language Model, Prefetching

To read the full article, click here

×
Semiconductor IP