Vendor: Expedera Category: NPU

Neural engine IP - Balanced Performance for AI Inference

On-device AI is a must-have for many new designs.

Overview

On-device AI is a must-have for many new designs. Silicon architects look for solutions that support the latest AI technologies, like transformers and stable diffusion, while balancing performance and low power consumption with minimal latency.

The Origin™ E2 is a family of power and area optimized NPU IP cores designed for devices like smartphones and edge nodes. It supports video—with resolutions up to 4K and beyond— audio, and text-based neural networks, including public, custom, and proprietary networks.

Innovative Architecture

The Origin E2 neural engine uses Expedera’s unique packet-based architecture, which is far more efficient than common layer-based architectures. The architecture enables parallel execution across multiple layers achieving better resource utilization and deterministic performance. It also eliminates the need for hardware-specific optimizations, allowing customers to run their trained neural networks unchanged without reducing model accuracy. This innovative approach greatly increases performance while lowering power, area, and latency.

Specifications

Compute Capacity 0.5K to 10K INT8 MACs
Multi-tasking Run Multiple Simultaneous Jobs
Power Efficiency 18 TOPS/W effective; no pruning, sparsity or compression required (though supported)
Example Networks Supported ResNet, MobileNet, MobileNet SSD Inception V3, RNN-T, BERT, EfficientNet, FSR CNN, CPN, CenterNet, Unet, YOLO V3, YOLO V5, ShuffleNet2, others
Example Performance MobileNet V1 (226 x 226): 8750 IPS, 13,482 IPS/W (N7 process, 1GHz, no sparsity/pruning/compression applied)
Layer Support Standard NN functions, including Conv, Deconv, FC, Activations, Reshape, Concat, Elementwise, Pooling, Softmax, others. Programmable general FP function, including Sigmoid, Tanh, Sine, Cosine, Exp, others, custom operators supported.
Data types INT4/INT8/INT10/INT12/INT16 Activations/Weights
FP16/BFloat16 Activations/Weights
Quantization Channel-wise Quantization (TFLite Specification)
Software toolchain supports Expedera, customer-supplied, or third-party quantization
Latency Deterministic performance guarantees, no back pressure
Frameworks TensorFlow, TFlite, ONNX, others supported

Key features

  • Choose the Features You Need: Customization brings many advantages, including increased performance, lower latency, reduced power consumption, and eliminating dark silicon waste. Expedera works with customers to understand their use case(s), PPA goals, and deployment needs during their design stage. Using this information, we configure Origin IP to create a customized solution that perfectly fits the application.
  • Market-Leading 18 TOPS/W: Sustained power efficiency is key to successful AI deployments. Continually cited as one of the most power-efficient architectures in the market, Origin NPU IP achieves a market-leading, sustained 18 TOPS/W.
  • Efficient Resource Utilization: Origin IP scales from GOPS to 128 TOPS in a single core. The architecture eliminates the memory sharing, security, and area penalty issues faced by lower-performing, tiled AI accelerator engines. Origin NPUs achieve sustained utilization averaging 80%—compared to the 20-40% industry norm—avoiding dark silicon waste.
  • Full TVM-Based Software Stack: Origin uses a TVM-based full software stack. TVM is widely trusted and used by OEMs worldwide. This easy-to-use software allows the importing of trained networks and provides various quantization options, automatic completion, compilation, estimator and profiling tools. It also supports multi-job APIs.
  • Successfully Deployed in 10M Devices: Quality is key to any successful product. Origin IP has successfully deployed in over 10 million consumer devices, with designs in multiple leading-edge nodes.

Block Diagram

Benefits

  • 1-20 TOPS performance
  • Support for standard, custom, and proprietary neural networks
  • Performance efficiencies up to 18 TOPS/Watt
  • Full software stack provided, including compiler, estimator, scheduler, and quantizer
  • Runs LLM, CNN, RNN, DNN, LSTM, and other network types
  • Delivered as Soft IP (RTL) or GDS

Specifications

Identity

Part Number
Origin E2
Vendor
Expedera
Type
Silicon IP

Files

Note: some files may require an NDA depending on provider policy.

Provider

HQ: USA

Learn more about NPU IP core

How To Start Building Edge-Native AI

Edge-native AI shifts inference to devices such as phones, cars, and sensors, enabling real-time processing, enhanced privacy, and operational resilience.

Why Vision LLMs Force A Rethink Of Edge AI Hardware

As vision-centric large language models move on-device, performance measured in raw TOPS is no longer enough. Architectures need to be built around real workloads, memory behavior, and sustained utilization, especially at the edge.

Benchmarking an NPU at Scale

With every SDK release, Quadric automatically recompiles and re-profiles the entire model zoo across a broad sweep of hardware configurations. The results land in DevStudio for any SoC architect to explore — no sales call, no handpicked numbers.

Your NPU Learned to Listen

Whisper runs end-to-end on Chimera. Encoder and decoder compiled as native GPNPU kernels, INT4 weights, FP16 attention, top-1 token match against the float32 reference. Scales to four cores with a flag, no recompile.

Heterogeneous NPU Data Movement Tax: Intel's Own Slides Tell the Story

At Quadric, we have long argued that heterogeneous NPU designs — those that stitch together multiple specialized fixed-function engines — carry an unavoidable hidden cost: data has to move. A lot. And data movement burns power, adds latency, and creates silicon-area overhead that scales with every new generation of AI models. Now, Intel has made that case for us.

Frequently asked questions about NPU IP cores

What is Neural engine IP - Balanced Performance for AI Inference?

Neural engine IP - Balanced Performance for AI Inference is a NPU IP core from Expedera listed on Semi IP Hub.

How should engineers evaluate this NPU?

Engineers should review the overview, key features, supported foundries and nodes, maturity, deliverables, and provider information before shortlisting this NPU IP.

Can this semiconductor IP be compared with similar products?

Yes. Buyers can compare this product with similar semiconductor IP cores or IP families based on category, provider, process options, and structured technical specifications.

×
Semiconductor IP