Terracotta: Enabling the Adoption of New DRAM Techniques via a Flexible DRAM Interface and Memory Controller

By Harsh Songara 1, Konstantinos Kanellopoulos 1, F. Nisa Bostancı 1, Konstantinos Marios Sgouras 1, Ataberk Olgun 1, İsmail Emir Yüksel 1, Andreas Kosmas Kakolyris 1, A. Giray Yağlıkçı 2, Onur Mutlu 1,3
1 ETH Zürich
2 CISPA 
3 New York University

Abstract

DRAM is the dominant memory technology in modern systems, yet it continues to limit performance, energy efficiency, and system robustness. To mitigate such limitations, many prior works have proposed novel DRAM techniques that change how DRAM is operated to support in-DRAM computation, improve memory access latency and parallelism, and enhance DRAM maintenance and reliability. However, adopting these techniques one after another in current systems with rigid interfaces and memory controller hardware requires repeated and costly modifications to them, hindering the deployment of critical performance, energy efficiency, and robustness improvements.

Our goal is to reduce the repeated DRAM interface and memory controller modifications that hinder or slow down the deployment of diverse DRAM techniques. We observe that the DRAM commands and memory controller structures across several DRAM techniques are similar. Our key idea is to leverage these similarities to compose a set of primitives that DRAM vendors and system designers can use to implement a wide variety of DRAM techniques. To this end, we propose Terracotta, a new framework that consists of two flexible components: (i) custom command extensions that enable DRAM vendors to define new commands within a single, standardized interface, and (ii) a programmable memory controller that system designers can easily and flexibly program to support new DRAM techniques post-silicon fabrication. Together, these two components enable the deployment of new DRAM techniques via configuration of the memory controller rather than repeated interface and controller modifications.

We design Terracotta for a DDR5-based system and comprehensively evaluate its performance, energy, and hardware complexity. We draw three key findings. First, for four DRAM techniques from four distinct domains including processing-using-DRAM, low-cost DRAM maintenance, subarray-level parallelism, and latency reduction, Terracotta retains almost all of the performance benefits (>96%) that custom implementations provide over a baseline system without these techniques. Second, a Terracotta based composition of two DRAM techniques from the domains of subarray-level parallelism and latency reduction outperforms the Terracotta-based implementations of the individual techniques, showing that Terracotta can enable system improvements by adding techniques without further interface and con troller modifications. Third, Terracotta’s flexibility incurs low DRAM energy overheads (0.6–3.2%) and low area and power overheads (0.03% and 0.56%, respectively) in a high-end server grade processor. Terracotta’s source code is freely available at https://github.com/CMU-SAFARI/Terracotta.

To read the full article, click here

×
Semiconductor IP