Reducing Instruction-Fetch Energy in RISC-V for Embedded AI Processing via Dynamic and Static Loop Caching
By Wiebren Wijnstra, Sameed Sohail, Berend-Jan van der Zwaag, Sabih Gerez, Amirreza Yousefzadeh
University of Twente, Enschede, The Netherlands

Abstract
Embedded RISC-V processors are increasingly deployed for on-device AI inference at the edge, where energy efficiency is a primary design constraint. Instruction fetching from SRAM-based memory is a dominant source of energy consumption in these cores, accounting for over 40% of total energy in our baseline measurements. This paper presents two loop cache architectures integrated into the datapath of a RISC-V processor: a dynamic loop cache that automatically detects and caches short backward-branch loops at runtime, and a static loop cache that functions as a software-managed hot-code instruction buffer, allowing preloading of arbitrary instruction blocks during the boot sequence. Both designs are implemented in the open-source NEORV32 RISC-V processor and evaluated on a LeNet-5 convolutional neural network inference workload, synthesized on GlobalFoundries 22nm FDX+ technology at 0.5V and 250MHz. The dynamic cache reduces instruction fetches by 48.3% and total energy by 21.5%, while the static cache achieves an 83.3% fetch reduction and 35.5% total energy savings. The area overhead of both designs remains below 0.2% of the full SoC area. The complete implementation is open source: https://github.com/wwiebren/NEORV32_loopcache_for_SparkRV.
Index Terms — RISC-V, loop cache, energy efficiency, instruction fetch, embedded AI processor, low-power edge inference
To read the full article, click here
Related Semiconductor IP
- RISC-V Debug & Trace IP
- RISC-V IOPMP IP
- Gen#2 of 64-bit RISC-V core with out-of-order pipeline based complex
- 64-bit RISC-V core with in-order single issue pipeline. Tiny Linux-capable processor for IoT applications.
- Multi-core capable RISC-V processor with vector extensions
Related Articles
- RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
- RISC-V in 2025: Progress, Challenges,and What’s Next for Automotive & OpenHardware
- e-GPU: An Open-Source and Configurable RISC-V Graphic Processing Unit for TinyAI Applications
- Boosting RISC-V SoC performance for AI and ML applications
Latest Articles
- Reducing Instruction-Fetch Energy in RISC-V for Embedded AI Processing via Dynamic and Static Loop Caching
- SPARC: Automated Root-Cause Analysis of Pre-Silicon Power Side-Channel Leakage in the Processor Design Flow
- A Heterogeneous Neural Network Accelerator for End-to-End Multitask RF Signal Recognition
- Hardware-Software Co-Design for Float16 On-Device Training on RISC-V Single-Core
- Low-Energy Reduced RISC-V Instruction Subset Processor for Tsetlin Machine Inference at the Edge