REACH: Controller-Managed Long-Span ECC for HBM AI Inference
By Rui Xie 1, Yunhua Fang 1, Asad Ul Haq 1, Linsen Ma 1, Sanchari Sen 2, Swagath Venkataramani 2, Liu Liu 1, Tong Zhang 1
1 Rensselaer Polytechnic Institute, Troy, NY 12180 USA
2 IBM T.J. Watson Research Center, Yorktown Heights, NY 10598 USA

Abstract
High-Bandwidth Memory (HBM) cost motivates stronger controller protection that can support a wider range of device error rates. Long-span error-correcting codes provide stronger protection at a comparable code rate, but a direct implementation couples small accesses to span-wide state and requires costly decoding at HBM bandwidth. Read-dominated LLM decode offers a favorable setting: sequential reads support span aggregation, while sparse writes limit parity-update traffic. This paper presents REACH, a controller microarchitecture that uses established inner codes to correct common errors and identify unresolved chunks, reserving a long outer code for known-erasure repair. Differential parity bounds write traffic, and a co-designed endpoint preserves 32 B transactions without an extra data burst. Ramulator2 sustains 1.88\,TB/s of application traffic at the highest error stress, while separate full-interface sizing supports a 2.69 TB/s application target using ASAP7-synthesized kernels. At this analytical target, REACH's nominal composition uses 55.8\% less controller area and 57.7% less modeled power than the evaluated mean-work direct-long design, showing the benefit of reserving long-span recovery for exceptional requests.
Index Terms—High-bandwidth memory (HBM), error correcting code (ECC), memory controller microarchitecture, Reed–Solomon codes, concatenated coding, erasure decoding, reliability, AI inference.
To read the full article, click here
Related Semiconductor IP
- HBM Memory Model
- HBM DFI Verification IP
- HBM 4 Verification IP
- HBM Synthesizable Transactor
- HBM DFI Synthesizable Transactor
Related Articles
- Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
- Making Strong Error-Correcting Codes Work Effectively for HBM in AI Inference
- AIA: A 16nm Multicore SoC for Approximate Inference Acceleration Exploiting Non-normalized Knuth-Yao Sampling and Inter-Core Register Sharing
- Why Software is Critical for AI Inference Accelerators
Latest Articles
- REACH: Controller-Managed Long-Span ECC for HBM AI Inference
- FPGA Acceleration of Fully Homomorphic Encryption with Adaptive Key Switching
- AutoTrans: AI-Assisted Automatic Translation of Security Assertions for RISC-V Processors
- LLM Inference on IMC-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism
- AI-Assisted Design of a Post-Quantum Cryptographic Accelerator: A Deployed-Silicon Case Study