Vendor: Tenstorrent Category: NPU

Tensix Cores for specialized acceleration

Tensix NEO® is Tenstorrent’s latest high performance AI architecture, optimized for performance-per-mm2 and performance-per-watt.

Overview

Tensix NEO® is Tenstorrent’s latest high performance AI architecture, optimized for performance-per-mm2 and performance-per-watt. The main unit of compute is the Tensix NEO Cluster, comprised of four Tensix NEO Cores, 4MB of shared L1 SRAM, high-bandwidth on-chip network, and data movement cores handling core-to-core, DRAM, and Ethernet I/O. Micro architectural optimizations maximize FPU utilization and performance in the Tensix NEO Cores while updating the ISA to support more complex data patterns and modern data formats like the MX formats and FP4.

Toolchain and Support

  • Open source simulation models available
  • Python and C++ entry points for operators, with paths for kernel customization and optimization
  • Powered by open source software:
  • The TT-Forge™ AI compiler allows for rapid model bring up and automatic optimization
  • The TT-Metalium™ SDK provides access directly to the metal, enabling use of Python and C++ for a variety of workloads
  • The TT-LLK SDK allows for tensor level programming

Key features

  • Updated mesh architecture that is both power-efficient and user friendly
  • Performant data movement from custom-tailored data movement cores
  • Improved performance-per-watt and performance per mm2; each Tensix NEO cluster contains:
    • 4 Tensix NEO Cores
    • Shared 4MB L1 SRAM
    • Shared NoC and Data Movement Block
  • Support for FP4 and MX data formats
  • Parallel issue of vector and matrix instructions
  • ASIL-B certified (NEO Auto only)

Block Diagram

Specifications

Identity

Part Number
Tensix Neo
Vendor
Tenstorrent
Type
Silicon IP

Files

Note: some files may require an NDA depending on provider policy.

Provider

HQ: Canada

Learn more about NPU IP core

Benchmarking an NPU at Scale

With every SDK release, Quadric automatically recompiles and re-profiles the entire model zoo across a broad sweep of hardware configurations. The results land in DevStudio for any SoC architect to explore — no sales call, no handpicked numbers.

Your NPU Learned to Listen

Whisper runs end-to-end on Chimera. Encoder and decoder compiled as native GPNPU kernels, INT4 weights, FP16 attention, top-1 token match against the float32 reference. Scales to four cores with a flag, no recompile.

Heterogeneous NPU Data Movement Tax: Intel's Own Slides Tell the Story

At Quadric, we have long argued that heterogeneous NPU designs — those that stitch together multiple specialized fixed-function engines — carry an unavoidable hidden cost: data has to move. A lot. And data movement burns power, adds latency, and creates silicon-area overhead that scales with every new generation of AI models. Now, Intel has made that case for us.

The Upcoming NPU Shakeout

The IP industry is no stranger to boom and bust cycles, and it looks to be at the crest of another wave.

Frequently asked questions about NPU IP cores

What is Tensix Cores for specialized acceleration?

Tensix Cores for specialized acceleration is a NPU IP core from Tenstorrent listed on Semi IP Hub.

How should engineers evaluate this NPU?

Engineers should review the overview, key features, supported foundries and nodes, maturity, deliverables, and provider information before shortlisting this NPU IP.

Can this semiconductor IP be compared with similar products?

Yes. Buyers can compare this product with similar semiconductor IP cores or IP families based on category, provider, process options, and structured technical specifications.

×
Semiconductor IP