A Framework for Accelerating Transformer Inference on RISC-V for Edge AI

By Ajay Kumar M 1, Vishnu PS 1, Yike Li 1, Shreejith Shanker 2, Dimitrios S. Nikolopoulos 3, Bo Ji 3, Hans Vandierendonck 4, Deepu John 1
1 University College Dublin, Ireland 
2 Trinity College Dublin, Ireland
3 Virginia Tech, USA 
4 Queen’s University Belfast, UK

Abstract

This work presents a framework for accelerating transformer-based language models (LMs) on resource-constrained IoT devices. The framework targets compact LMs: BERT-Tiny (B-Ty), MobileBERT (M-Bt), MiniLM (M-Lm), Electra (E-Lt) and DeBERTa (D-Bt) -- selected for their architectural diversity and use in edge inference scenarios. The proposed flow derives lightweight instruction set extensions tailored to the non-obvious computational patterns of these models. In addition, a custom instruction is introduced to accelerate the address generation stage of batch matrix multiplication, achieving a performance improvement of 15.39--21.74% with modest ASIC overheads of 6.79% in area and 2.33% in power. To further enhance performance without incurring additional processor core hardware cost, an optional compiler-directed loop unrolling strategy is employed, trading increased code size for overall reduced execution time. Evaluation on the Synopsys trv32p3f RISC-V core demonstrates inference speedups of up to 2.19x, and FPGA implementation on the AMD Zynq UltraScale+ ZCU102 shows a 32.94% area overhead at 75 MHz, whereas the ASIC implementation using the TSMC 28 nm library incurs a 21.08% area overhead while operating at 250 MHz.

Index Terms—RISC-V, Language models, TVM, Hardware acceleration, FPGA

To read the full article, click here

×
Semiconductor IP