Researchers Explore Next-Generation AI Architectures Using Weebit ReRAM

As artificial intelligence expands beyond data centers and into edge devices, autonomous systems, industrial equipment, and even spacecraft, power efficiency is becoming one of the industry’s most important challenges.

While AI models continue to grow in capability, conventional computing architectures still rely on constant movement of data between memory and processors. This data movement consumes energy, increases latency, and creates system-level bottlenecks, particularly in environments where power and thermal budgets are limited.

To address this challenge, researchers are increasingly investigating In-Memory Computing (IMC) architectures that bring computation closer to where data resides. By performing computation within memory arrays, such architectures can significantly improve efficiency while reducing latency and system complexity.

ReRAM is particularly well suited to IMC research because the same memory cells used to store neural-network weights can also participate in analog computation within crossbar arrays. This reduces data movement while allowing many multiply-accumulate operations to occur in parallel, making the architecture attractive for energy-constrained AI applications.

Recently, a U.S. based company developing radiation-tolerant AI hardware used Weebit ReRAM to develop and evaluate a next-generation IMC AI accelerator. The work was part of a project focused on developing low-power AI inference capabilities for harsh-environment applications.

The project explored neural-network operations implemented directly on ReRAM-based compute arrays. Researchers were able to characterize memory behavior, evaluate compute-in-memory techniques, and benchmark the architecture against GPU- and FPGA-based implementations of the same inference workload.

Above: Results from the project per given AI workload*

The project highlighted one of the key advantages of IMC: reducing the energy associated with moving data between memory and processing units. Performance projections for the benchmarked image classification workload, based on a custom convolutional neural network (CNN) implementation, indicate that the Weebit ReRAM-based IMC array would consume approximately 13x less power than the FPGA-based implementation and nearly 90x less power than the GPU-based implementation.


Explore Weebit IP:


The projections also indicate that the implementation would reduce inference latency from tens of milliseconds to below one millisecond and achieve approximately 280x higher compute efficiency (GOPS/W) than the FPGA-based accelerator.

While embedded ReRAM is already being commercialized as a replacement for embedded flash, research such as this demonstrates its potential to support future memory-centric computing architectures. As the industry seeks more efficient ways to execute increasingly demanding AI workloads, advanced non-volatile memories such as ReRAM (RRAM) are attracting growing interest in applications where power efficiency and performance are critical.

Although much work remains before IMC becomes mainstream, projects like this illustrate the direction of AI hardware research and the important role memory technologies are expected to play.

Learn more about how Weebit ReRAM can enable new approaches that overcome the limitations of traditional computing architectures.

 

*GPU: NVIDIA GeForce RTX 4090; FPGA: Xilinx ZCU104 (Zynq UltraScale+); IMC system-level performance projected from measured Weebit ReRAM-based array using application-level modeling.

×
Semiconductor IP