FlexViT: A Flexible FPGA-based Accelerator for Edge Vision Transformers
By Hubert Dymarkowski, Xingjian Fu, Rappy Saha, Jude Haris, José Cano
University of Glasgow, Scotland, UK

Abstract
Deploying Vision Transformer (ViT) models on edge platforms remains challenging due to their high computational demands and the architectural heterogeneity of modern hybrid ViT models, which incorporate both fully connected and convolutional layers. This heterogeneity leads to significant variation in tensor shapes, requiring flexible and efficient FPGA-based acceleration. In this paper, we present FlexViT, a reconfigurable FPGA accelerator for efficient ViT inference on resource-constrained edge devices. Built on the SECDA-TFLite framework, FlexViT employs a hardware-software co-design approach that maps both fully connected and convolutional layers onto a unified high-throughput INT8 GEMM engine using a runtime im2col transformation. To efficiently support diverse layer configurations, we propose a dual-mode dataflow that dynamically switches between input and weight reuse by reconfiguring the compute array at runtime. We further introduce a depth-first tiling strategy that completes accumulation in a single pass, eliminating off-chip partial-sum transfers and reducing memory bandwidth requirements. We implement FlexViT on a PYNQ-Z2 FPGA and evaluate it across a representative set of ViT models. FlexViT achieves up to 2.74x speedup on accelerator-executed layers, translating into up to 1.40x end-to-end speedup compared to CPU-only execution. The code is available at: https://github.com/gicLAB/FlexViT
Index Terms—Vision Transformers, FPGA, Edge AI.
To read the full article, click here
Related Semiconductor IP
- NPU IP
- JPEG XL Encoder
- I2C Master/Slave Controller Core
- NVMe Validation Test Suite
- Hybrid Memory Cube Verification IP
Related Articles
- VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
- A Flexible Sparsity-Aware FPGA Accelerator with Column-Wise Compression for Efficient CNN Inference
- BitFair: A 12nm Bit-Serial CNN Accelerator with Learnable Early Termination and Adaptive Bit Ordering for Ultra-Low-Power XR Vision
- Peregrino: A Full-Hardware Accelerator for the Complete Falcon Post-Quantum Digital Signature Scheme on Resource-Constrained Edge Devices
Latest Articles
- Terracotta: Enabling the Adoption of New DRAM Techniques via a Flexible DRAM Interface and Memory Controller
- A Framework for Accelerating Transformer Inference on RISC-V for Edge AI
- An Interleaved Parallel Dependent Quantization Hardware Architecture for H.266/VVC
- A Formal Security Analysis of CAN XL
- A Secure dToF LiDAR SoC with Dual-Domain Fingerprinting and Event-Driven AFE Circuit Achieving Sensor-Level Attack Resilience