Execute-in-Place: Getting More from Embedded NVM

When designing an embedded system, the location where code is stored is only part of the memory equation. Designers also need to consider how that code gets to the processor when it is time to execute it.

In many embedded architectures based on a microcontroller (MCU), program code is stored in non-volatile memory (NVM) and copied into SRAM (either as tightly-coupled memory or cache memory) for execution. This approach provides fast access to code, but it also requires SRAM capacity for code storage, and it takes time and energy to move code from one memory to another.

Execute-in-Place, or XiP, offers another approach: the processor executes code directly from the NVM where it is stored, without first copying it into SRAM.

XiP itself is not a new concept. But embedded systems are becoming more capable while facing tighter power and area constraints. In this context, fast embedded NVM technologies such as ReRAM make it worth taking a fresh look at XiP and the system-level benefits it can provide.

 

Reducing data movement and SRAM requirements

With XiP, code remains in NVM and the processor fetches instructions directly from it, avoiding the need to first copy the code into SRAM. This can reduce data movement and accelerate system wake-up, particularly in systems that spend much of their time in a low-power state and periodically wake to perform a task.

XiP can also reduce the SRAM capacity needed for code storage. This can save silicon area and also reduce SRAM leakage power, an important consideration in battery-operated IoT and medical devices that spend extended periods in low-power modes.


Explore Weebit Nano IP:


Why ReRAM is well suited for XiP

For XiP to be practical, the NVM must be able to deliver code to the processor quickly enough for the target application.

Weebit ReRAM combines non-volatility with fast, low-power read access. As we discussed in a previous article on Weebit ReRAM power consumption, read operations are the most frequent NVM operation, making read performance and power particularly important in embedded systems. Weebit ReRAM reads from a low-voltage power supply and does not require an always-on charge pump.

Many embedded applications operate at relatively modest MCU frequencies, making ReRAM’s fast read performance well suited to XiP. Typical MCU vendors implement XiP at around 40-50 MHz in SoCs where this operating frequency meets the application requirements.

Weebit’s ReRAM module supports a dedicated XiP interface, with pipelined reads and operation at up to 100 MHz, corresponding to a 10ns read cycle. This allows the processor to access code directly from the ReRAM module without first transferring it to SRAM.

ReRAM can also provide important area advantages. A ReRAM bitcell can be 4-6x smaller than an SRAM bitcell, helping reduce the silicon area required for code storage. This complements the potential area savings from reducing the amount of SRAM needed to hold code for execution.

ReRAM also offers advantages as semiconductor designs move to smaller process geometries. Unlike embedded flash, ReRAM can scale to advanced process nodes. And because it is integrated in the back end of line (BEOL), it can be added with relatively few additional process steps and masks.

More than read-only code storage

Storing code directly in ReRAM can provide another advantage beyond XiP: the ability to update that code efficiently. Weebit ReRAM combines low programming energy with bit-level accessibility, enabling targeted code updates without the need to erase and rewrite large blocks of memory.

This can make it easier to support code patches, bug fixes and new features after deployment. It can also be particularly valuable for firmware-over-the-air (FOTA) updates in battery-operated devices, where the energy required to update non-volatile memory is an important system consideration.

What can XiP mean at the system level?

The benefits become clearer when XiP is considered as part of the complete system architecture.

In an earlier Weebit analysis of memory partitioning in an IoT application, we considered a wearable sensor that periodically wakes to log and process data. In a conventional architecture, code is stored in external flash and loaded into local code SRAM when the system wakes.

We then considered an alternative architecture in which on-chip ReRAM (RRAM) stores the code and the MCU fetches it directly from ReRAM using XiP. This removes the need to access the external flash during wake-up and allows the code ReRAM to be powered down when it is not required. Our calculations for this specific use case showed a 30% power saving compared with the traditional architecture.

The same analysis looked at additional memory partitioning options, including storing logged data in ReRAM and replacing portions of data SRAM with ReRAM. This illustrates a broader point: selecting an embedded NVM is not simply a question of replacing one memory technology with another. Its characteristics can give designers new options for partitioning memory across the system.

Looking beyond NVM as storage

Non-volatile memory is often thought of primarily as a place to retain code and data when power is removed. ReRAM opens up a broader role. Its fast read performance can enable direct code execution through XiP, while its density, scalability, and low-energy programming can provide additional flexibility for how code is stored and updated. For embedded-system designers, that means the choice of NVM can influence not only how information is stored, but also the architecture, power consumption, and the capabilities of the overall system.

×
Semiconductor IP