A 129FPS Full HD Real-Time Accelerator for 3D Gaussian Splatting
By Fang-Chi Chang, and Tian-Sheuan Chang
Institute of Electronics, National Yang Ming Chiao Tung University, Taiwan

Abstract
Rendering large-scale, unbounded scenes on AR/VR-class devices is constrained by the computation, bandwidth, and storage cost of 3D Gaussian Splatting (3DGS). We propose a low-power, low-cost 3DGS hardware accelerator that renders full-HD images in real time, together with a hardware-friendly compression pipeline that combines iterative Gaussian pruning and fine-tuning, progressive spherical harmonics (SH) degree reduction, and vector quantization of all SH coefficients and colors. The scheme achieves a 51.6× model-size reduction with a 0.743 dB PSNR loss. The accelerator uses a frame-level pipeline that integrates point-based culling and projection with tile-based sorting and rasterization, skips zero-Jacobian matrix multiplications (reducing processing elements by 63\% and computation by 53\%), and adopts comparison-free tile-based sorting with deterministic latency. Implemented in a TSMC 28-nm process at 800 MHz, the design occupies 0.66 mm2 with 1.1438 M gates and 120 kB SRAM, consumes 0.219 W, and delivers 1219 Mpixels/J at 267.5 Mpixels/s, enabling 1080p at 129 FPS. Overall, it is 5.98× smaller in area, 5.94× higher throughput, and delivers 7.5× higher energy efficiency than prior 3DGS accelerators.
Index Terms—3D Gaussian Splatting, hardware accelerators, model compression, real-time rendering
To read the full article, click here
Related Semiconductor IP
- nQrux® Root of Trust IP
- AXI to UCIe Bridge IP
- UCIe 2.x Controller IP
- SWI3S (SoundWire I3S Interface) Peripheral Controller Core IP
- OpenTitan-based RISC-V Secure Element
Related Articles
- Peregrino: A Full-Hardware Accelerator for the Complete Falcon Post-Quantum Digital Signature Scheme on Resource-Constrained Edge Devices
- Vorion: A RISC-V GPU with Hardware-Accelerated 3D Gaussian Rendering and Training
- A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
- A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
Latest Articles
- Automated Pre-Silicon Verification of High-Speed DDR5 and LPDDR5/6 Memory Controllers: Closed-Loop Timing, Mode Register, and PHY Synchronization in UVM
- U-Sonic: An Open-Source 8-Channel Ultrasound Transmit IP in a 130 nm RISC-V SoC
- S-ALSA: Co-Design of Adiabatic Logic-based Sensing and Balanced Bit-Cells for Secure and Energy-Efficient MRAM
- MEGATRON: a 28nm Analog PCM CiM/Digital System-on-Chip for Edge GenAI at 57.5 TOPS/W and 1.52 Mparam/mm²
- Peregrino: A Full-Hardware Accelerator for the Complete Falcon Post-Quantum Digital Signature Scheme on Resource-Constrained Edge Devices