Physical Design Exploration of a Wire-Friendly Domain-Specific Processor for Angstrom-Era Nodes
By Lorenzo Ruotolo 1, 4, Lara Orlandic 2, Pengbo Yu 2, Moritz Brunion 4, Daniele Jahier Pagliari 1, Dwaipayan Biswas 4, Giovanni Ansaloni 2, David Atienza 2, Julien Ryckaert 4, Francky Catthoor 3, and Yukai Chen 4
1 Politecnico di Torino, Italy
2 Ecole Polytechnique Fédérale de Lausanne (EPFL), Switzerland
3 National Technical University of Athens, Greece
4 IMEC, Belgium

Abstract
This paper presents the physical design exploration of a domain-specific processor (DSIP) architecture targeted at machine learning (ML), addressing the challenges of interconnect efficiency in advanced Angstrom-era technologies. The design emphasizes reduced wire length and high core density by utilizing specialized memory structures and SIMD (Single Instruction, Multiple Data) units. Five configurations are synthesized and evaluated using the IMEC A10 nanosheet node PDK. Key physical design metrics are compared across configurations and against VWR2A, a state-of-the-art (SoA) DSIP baseline. Results show that our architecture achieves over 2x lower normalized wire length and more than 3x higher density than the SoA, with low variability in the metrics across all configurations, making it a promising solution for next-generation DSIP designs. These improvements are achieved with minimal manual layout intervention, demonstrating the architecture's intrinsic physical efficiency and potential for low-cost wire-friendly implementation.
Index Terms—domain-specific processors, machine learning, physical design, nanosheet, wirelength optimization.
To read the full article, click here
Related Semiconductor IP
- Neuromorphic Processor IP (Second Generation)
- Processor Development Toolset
- Low Power 32-bit Processor
- Nios® V Processor
- ARC-V RHX-100 dual-issue, 32-bit single-core RISC-V processor for real-time applications
Related Articles
- Customizing a Large Language Model for VHDL Design of High-Performance Microprocessors
- TensorPool: A 3D-Stacked 8.4TFLOPS/4.3W Many-Core Domain-Specific Processor for AI-Native Radio Access Networks
- Understanding the Importance of Prerequisites in the VLSI Physical Design Stage
- Inside HDR10: A technical exploration of High Dynamic Range
Latest Articles
- New Number Formats for FFT IP Cores in Optical OFDM Transceivers
- Reducing Power Consumption of Embedded Dynamic Memories with ECCs
- NIFA: Nonlinear IMC enhanced FPGA for efficient ML inference
- A 32-channel event-based bio-signal analog front-end with adaptive delta and pulse frequency encoding
- Vectorizing Quantum Control: A RISC-V Vector Extension Architecture for Scalable Qubit Systems