REACH: Controller-Managed Long-Span ECC for HBM AI Inference

By Rui Xie 1, Yunhua Fang 1, Asad Ul Haq 1, Linsen Ma 1, Sanchari Sen 2, Swagath Venkataramani 2, Liu Liu 1, Tong Zhang 1
1 Rensselaer Polytechnic Institute, Troy, NY 12180 USA
2 IBM T.J. Watson Research Center, Yorktown Heights, NY 10598 USA

Abstract

High-Bandwidth Memory (HBM) cost motivates stronger controller protection that can support a wider range of device error rates. Long-span error-correcting codes provide stronger protection at a comparable code rate, but a direct implementation couples small accesses to span-wide state and requires costly decoding at HBM bandwidth. Read-dominated LLM decode offers a favorable setting: sequential reads support span aggregation, while sparse writes limit parity-update traffic. This paper presents REACH, a controller microarchitecture that uses established inner codes to correct common errors and identify unresolved chunks, reserving a long outer code for known-erasure repair. Differential parity bounds write traffic, and a co-designed endpoint preserves 32 B transactions without an extra data burst. Ramulator2 sustains 1.88\,TB/s of application traffic at the highest error stress, while separate full-interface sizing supports a 2.69 TB/s application target using ASAP7-synthesized kernels. At this analytical target, REACH's nominal composition uses 55.8\% less controller area and 57.7% less modeled power than the evaluated mean-work direct-long design, showing the benefit of reserving long-span recovery for exceptional requests.

Index Terms—High-bandwidth memory (HBM), error correcting code (ECC), memory controller microarchitecture, Reed–Solomon codes, concatenated coding, erasure decoding, reliability, AI inference.

To read the full article, click here

×
Semiconductor IP