A Framework for Accelerating Transformer Inference on RISC-V for Edge AI
By Ajay Kumar M 1, Vishnu PS 1, Yike Li 1, Shreejith Shanker 2, Dimitrios S. Nikolopoulos 3, Bo Ji 3, Hans Vandierendonck 4, Deepu John 1
1 University College Dublin, Ireland
2 Trinity College Dublin, Ireland
3 Virginia Tech, USA
4 Queen’s University Belfast, UK

Abstract
This work presents a framework for accelerating transformer-based language models (LMs) on resource-constrained IoT devices. The framework targets compact LMs: BERT-Tiny (B-Ty), MobileBERT (M-Bt), MiniLM (M-Lm), Electra (E-Lt) and DeBERTa (D-Bt) -- selected for their architectural diversity and use in edge inference scenarios. The proposed flow derives lightweight instruction set extensions tailored to the non-obvious computational patterns of these models. In addition, a custom instruction is introduced to accelerate the address generation stage of batch matrix multiplication, achieving a performance improvement of 15.39--21.74% with modest ASIC overheads of 6.79% in area and 2.33% in power. To further enhance performance without incurring additional processor core hardware cost, an optional compiler-directed loop unrolling strategy is employed, trading increased code size for overall reduced execution time. Evaluation on the Synopsys trv32p3f RISC-V core demonstrates inference speedups of up to 2.19x, and FPGA implementation on the AMD Zynq UltraScale+ ZCU102 shows a 32.94% area overhead at 75 MHz, whereas the ASIC implementation using the TSMC 28 nm library incurs a 21.08% area overhead while operating at 250 MHz.
Index Terms—RISC-V, Language models, TVM, Hardware acceleration, FPGA
To read the full article, click here
Related Semiconductor IP
- RISC-V Debug & Trace IP
- RISC-V IOPMP IP
- Gen#2 of 64-bit RISC-V core with out-of-order pipeline based complex
- 64-bit RISC-V core with in-order single issue pipeline. Tiny Linux-capable processor for IoT applications.
- Multi-core capable RISC-V processor with vector extensions
Related Articles
- FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
- VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
- MultiVic: A Time-Predictable RISC-V Multi-Core Processor Optimized for Neural Network Inference
- RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
Latest Articles
- A Framework for Accelerating Transformer Inference on RISC-V for Edge AI
- An Interleaved Parallel Dependent Quantization Hardware Architecture for H.266/VVC
- A Formal Security Analysis of CAN XL
- A Secure dToF LiDAR SoC with Dual-Domain Fingerprinting and Event-Driven AFE Circuit Achieving Sensor-Level Attack Resilience
- ZTA-Q: an Open-source RISC-V Platform for Accurate Quantized CNN Inference