Overview
Single-core neural network accelerator offering from 0.5 to 4 TOPS Optimized for machine learning inference applications
The Cadence® Tensilica® NNA 110 accelerator incorporates a custom hardware accelerator engine (NNE) coupled with a Tensilica Vision P6 or P1 DSP. The specialized compute block inside the NNA 110 hardware leverages features like random sparsity, tensor compression / decompression to provide an overall best in-class embedded AI accelerator solution.
A single-core NNA 110 accelerator supports 256 to 2K MAC 8x8-bit MAC computations and has various user-defined configurable options. The NNA 110 accelerator can run all neural network layers, including but not limited to convolution, fully connected, LSTM, LRN, and pooling operations. The accompanying Tensilica DSP in NNA 110 can run any operation that is not native to the accelerator, thereby making NNA 110 a highly flexible and robust future-proof offering. NNA 110 solution deliverables comprises of turnkey soft RTL IP, software compiler toolchain, and an accurate simulator for benchmarking.
Learn more about NPU IP core
Passing PCI Express (PCIe) compliance is different from being ready for the field. A PCIe link can clear every test in a controlled lab environment and still develop margin problems six months into deployment.
Modern compute systems have evolved beyond reliance on a single dominant interface. Today, they're increasingly defined by their ability to support multiple high-speed protocols concurrently—including PCIe, Ethernet, and others. This shift toward multi-protocol capability is fundamentally reshaping how we architect intelligent edge AI systems, especially as inferencing workloads grow more distributed, data-intensive, and latency-sensitive.
The world of artificial intelligence is moving beyond the cloud and into our everyday devices from smart sensors to robotics and AR/VR headsets.
In the rapidly advancing field of artificial intelligence, two neural network architectures have become prominent: convolutional neural networks (CNNs) and transformers. Each architecture has brought significant advancements to various domains, ranging from image recognition, video surveillance to natural language processing (NLP), speech recognition and generation, multimodal AI and more. This article aims to compare their differences and respective strengths and highlight how specialized hardware like Tensilica Vision DSPs accelerates these models for real-world applications.
Artificial intelligence is rapidly expanding its reach into embedded systems and edge devices, driving the need for specialized processors that can efficiently handle complex AI workloads. While Neural Processing Units (NPUs) excel at accelerating neural network computations, the increasing demands of agentic and physical AI networks require a more comprehensive approach. This is where the Cadence Tensilica NeuroEdge 130 AI Co-Processor (AICP) comes in, designed to complement any NPU and enable end-to-end execution of these advanced AI applications.
Deploying PyTorch models on embedded devices, especially audio DSPs, presents unique challenges. To address these, Cadence and Meta have collaborated to create a robust, high-performance framework for deploying machine learning models on Cadence's Tensilica HiFi DSP family. By leveraging ExecuTorch and applying both graph-level and operator-level optimizations, the teams have achieved speedups of at least an order of magnitude compared to standard out-of-the-box deployments.