Physical Design Exploration of a Wire-Friendly Domain-Specific Processor for Angstrom-Era Nodes
By Lorenzo Ruotolo 1, 4, Lara Orlandic 2, Pengbo Yu 2, Moritz Brunion 4, Daniele Jahier Pagliari 1, Dwaipayan Biswas 4, Giovanni Ansaloni 2, David Atienza 2, Julien Ryckaert 4, Francky Catthoor 3, and Yukai Chen 4
1 Politecnico di Torino, Italy
2 Ecole Polytechnique Fédérale de Lausanne (EPFL), Switzerland
3 National Technical University of Athens, Greece
4 IMEC, Belgium

Abstract
This paper presents the physical design exploration of a domain-specific processor (DSIP) architecture targeted at machine learning (ML), addressing the challenges of interconnect efficiency in advanced Angstrom-era technologies. The design emphasizes reduced wire length and high core density by utilizing specialized memory structures and SIMD (Single Instruction, Multiple Data) units. Five configurations are synthesized and evaluated using the IMEC A10 nanosheet node PDK. Key physical design metrics are compared across configurations and against VWR2A, a state-of-the-art (SoA) DSIP baseline. Results show that our architecture achieves over 2x lower normalized wire length and more than 3x higher density than the SoA, with low variability in the metrics across all configurations, making it a promising solution for next-generation DSIP designs. These improvements are achieved with minimal manual layout intervention, demonstrating the architecture's intrinsic physical efficiency and potential for low-cost wire-friendly implementation.
Index Terms—domain-specific processors, machine learning, physical design, nanosheet, wirelength optimization.
To read the full article, click here
Related Semiconductor IP
- Neuromorphic Processor IP (Second Generation)
- Processor Development Toolset
- Low Power 32-bit Processor
- Nios® V Processor
- ARC-V RHX-100 dual-issue, 32-bit single-core RISC-V processor for real-time applications
Related Articles
- TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks
- Understanding the Importance of Prerequisites in the VLSI Physical Design Stage
- Inside HDR10: A technical exploration of High Dynamic Range
- The role of cache in AI processor design
Latest Articles
- A Low-Latency ASIC Architecture for Real-Time Line Segment Detection
- BitFair: A 12nm Bit-Serial CNN Accelerator with Learnable Early Termination and Adaptive Bit Ordering for Ultra-Low-Power XR Vision
- A Flexible Sparsity-Aware FPGA Accelerator with Column-Wise Compression for Efficient CNN Inference
- Reducing Instruction-Fetch Energy in RISC-V for Embedded AI Processing via Dynamic and Static Loop Caching
- SPARC: Automated Root-Cause Analysis of Pre-Silicon Power Side-Channel Leakage in the Processor Design Flow