面向FPGA上MLC NVM的保留感知RISC-V ISA扩展与内存控制器
Retention-Aware RISC-V ISA Extension and Memory Controller on FPGA for MLC NVM
AI总结:
该研究提出面向FPGA的MLC NVM内存控制器架构与RISC-V指令扩展,通过平衡写入速度和数据保留时间优化MLC NVM写入,实验显示其可降低硬件开销并提升性能。
AI中文摘要:
非易失性内存(NVM)技术,尤其是多层单元(MLC)NVM,在提升内存密度方面具有巨大潜力。MLC NVM在写入延迟和保留时间之间存在权衡:更快的写入/存储操作会导致保留时间降低,而更慢的写入则会带来更高的保留时间。然而,目前在硬件中验证和原型化基于NVM的系统、并在系统级利用这种权衡的工作十分有限。在本文中,我们提出了一种新颖的内存控制器架构和RISC-V指令集扩展,通过平衡速度和保留时间来优化MLC NVM的写入操作。我们的定制NVM控制器基于带有AXI内存映射接口的有限状态机构建,通过增强的突发传输高效管理读写操作,最大限度降低延迟。此外,我们在RISC-V中引入了fast-store指令,以提升写入性能,同时解决保留时间限制。我们还设计了一种专用AXI从设备外设,支持位感知写入:关键位(例如最高有效位MSB)使用较慢、高保留时间的写入,而非关键位(例如最低有效位LSB)使用较快、低保留时间的写入,从而在不损害数据可靠性的前提下提升性能。这些增强功能在FPGA平台上以硬件形式实现。实验结果显示,与传统设计相比,我们的控制器将硬件开销降低了30%;对于流式工作负载,fast-store指令将性能提升超过7%,硬件开销不到0.08%;位级AXI外设的查找表(LUT)利用率即使对于64×64矩阵也保持在3.5%以下,对于32×32大小则低于1%,使其可集成到更大的片上系统(SoC)中。
英文摘要:
Non-volatile memory (NVM) technologies, particularly Multi-Level Cell (MLC) NVMs, offer significant potential for increasing memory density. MLC NVMs provide a tradeoff between write latency and retention time, where faster writes/stores result in lower retention and slower writes yield higher retention. However, limited work has been done to validate and prototype NVM-based systems in hardware, leveraging this tradeoff at the system level. In this paper, we present a novel memory controller architecture and a RISC-V instruction set extension to optimize MLC NVM write operations by balancing speed and retention time. Our custom NVM controller, built around a finite state machine with an AXI memory-mapped interface, efficiently manages read/write operations with enhanced burst transfers, minimizing latency. Additionally, we introduce a fast-store instruction in RISC-V to increasing write performance while addressing retention limitations. Further, we design a dedicated AXI slave peripheral that supports bit-significance-aware writes: critical bits (e.g., MSBs) are written using slower, high-retention writes, while non-critical bits (e.g., LSBs) use faster, low-retention writes to help enhance performance without compromising data reliability. These enhancements are implemented in hardware on an FPGA platform. Experimental results show that our controller reduces hardware overhead by 30% compared to conventional designs, and the fast-store instruction improves performance by over 7% for streaming workloads with less than 0.08% hardware overhead. The bit-wise AXI peripheral has a LUT utilization staying below 3.5% even for 64x64 matrices, and under 1% for 32x32 sizes, making it viable for integration into larger SoCs.