NIFA: Nonlinear IMC enhanced FPGA for efficient ML inference
By Jiajun Hu 1, Ruthwik Reddy Sunketa 1, Lei Zhao 2, Archit Gajjar 2, Luca Buonanno 2, Aman Arora 1
1 Arizona State University, Tempe, AZ, USA
2 Hewlett Packard Enterprise Labs, Fort Collins, CO, USA

Abstract
Recent FPGAs have improved deep learning (DL) inference efficiency through dedicated tensor blocks and in-BRAM computation. ReRAM-based analog in-memory computing (IMC) pushes efficiency further, offering an order-of-magnitude improvement in compute density and energy efficiency over conventional digital logic by performing vector-matrix multiplication (VMM) directly within the ReRAM crossbar; prior work has integrated such IMC blocks into FPGAs for DL inference. However, conventional IMC designs support only static-weight VMM, leaving nonlinear operations and dynamic matrix-matrix multiplication (DIMM) to the FPGA fabric. As a result, the benefits of IMC are largely confined to static-weight models, whereas Transformer-based models, which rely on frequent nonlinear and DIMM operations, gain only limited improvement. Moreover, the ADCs within each IMC block consume more than 70% of its area and power, further limiting system efficiency and scalability. To address these limitations, we propose a novel FPGA architecture that integrates an ADC-free IMC block, replacing the conventional ADC with analog content-addressable memories (ACAMs) that natively perform nonlinear operations inside the block. To fully exploit this block, we conduct an FPGA-aware design-space exploration that determines optimal crossbar dimensions while balancing FPGA area, flexibility, and DL performance, and we develop an efficient mapping that leverages ACAMs to carry out DIMM operations, extending the applicability of IMC to attention computation. On CNN and Transformer-based benchmarks, the proposed architecture achieves up to 40x and 1.9x higher energy efficiency and 4.1x and 2.5x higher area efficiency, respectively. Overall, it significantly improves FPGA DL inference efficiency and sustains robust gains on Transformer-based workloads across long input sequences, advancing domain-specialized FPGA design.
To read the full article, click here
Related Semiconductor IP
- TSMC 4nm 3V3 GPIO
- 1 Kbyte EEPROM IP
- 3.6 Kbit EEPROM IP
- AMBA 5 AHB Bus TFT LCD / OLED Display Controller
- AMBA AXI5 / ACE5-Lite TFT LCD / OLED Display Controller
Related Articles
- A Flexible Sparsity-Aware FPGA Accelerator with Column-Wise Compression for Efficient CNN Inference
- Bare-Metal RISC-V + NVDLA SoC for Efficient Deep Learning Inference
- FPGA-Accelerated RISC-V ISA Extensions for Efficient Neural Network Inference on Edge Devices
- A Multiprocessor System-on-chip Architecture with Enhanced Compiler Support and Efficient Interconnect
Latest Articles
- BitFair: A 12nm Bit-Serial CNN Accelerator with Learnable Early Termination and Adaptive Bit Ordering for Ultra-Low-Power XR Vision
- A Flexible Sparsity-Aware FPGA Accelerator with Column-Wise Compression for Efficient CNN Inference
- Reducing Instruction-Fetch Energy in RISC-V for Embedded AI Processing via Dynamic and Static Loop Caching
- SPARC: Automated Root-Cause Analysis of Pre-Silicon Power Side-Channel Leakage in the Processor Design Flow
- A Heterogeneous Neural Network Accelerator for End-to-End Multitask RF Signal Recognition