arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PipeDRAM:一种无数据转置的DRAM内架构,支持硬件/软件流水线

PipeDRAM: A Data-Transposition-Free Processing-Using-DRAM Architecture with Hardware/Software Pipelining

Geraldo F. Oliveira, Ataberk Olgun, Ismail Emir Yüksel, F. Nisa Bostancı, Pedro H. E. Becker, Mohammad Sadrosadati, Saugata Ghose, Juan Gómez-Luna, Onur Mutlu

arXiv 2609.30998首次发表:更新:

发表机构

Huawei Research, Zürich; ETH Zürich; SAFARI Research Group(华为研究院苏黎世; 苏黎世联邦理工学院; SAFARI研究组)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PipeDRAM通过确定性位重组织和流水线执行模型,消除PUD架构中的运行时数据转置,实现水平数据布局上的直接位级运算,性能提升最高80.4倍,能耗降低最高38倍。

AI 中文摘要

利用DRAM进行处理(PUD)架构利用DRAM的模拟运算特性,通过将数据组织在垂直布局中(操作数位沿DRAM列堆叠),在存储阵列内部执行大规模位级布尔和算术运算。然而,现代计算系统原生采用水平数据布局,以保持缓存行抽象、利用行缓冲器中的空间局部性并实现高内存吞吐量。这种根本性不匹配迫使现有PUD架构频繁地在水平与垂直格式之间进行数据布局转换,导致显著的性能、能耗和系统集成开销。我们的目标是以低成本消除PUD系统中的数据转置开销。为此,我们提出PipeDRAM,一种PUD架构,消除了运行时数据布局转换的需求,使得PUD操作能够直接对水平布局的数据执行。PipeDRAM的关键思想是:(i)确定性地重新组织每个内存请求内部的位,以在水平数据布局中实现DRAM阵列内PUD友好的数据放置,以及(ii)采用基于流水线的执行模型,重叠位相关和位无关的DRAM内操作,以利用整个存储阵列的位级并行性。我们将PipeDRAM与不同的计算平台进行比较。与三种最先进的PUD系统相比,PipeDRAM提供了(i)11.8倍、11.8倍和80.4倍的更高性能,以及(ii)25.4倍、3.0倍和38.0倍的更低能耗。PipeDRAM在DRAM芯片(1.86%)和CPU芯片(0.05%)上引入了较低的面积开销。为了促进PUD系统的进一步研究,我们在以下https URL开源了PipeDRAM。

英文摘要

Processing-using-DRAM (PUD) architectures exploit the analog operational properties of DRAM to perform bulk bitwise Boolean and arithmetic operations inside memory arrays by organizing data in a vertical layout, where operand bits are stacked along DRAM columns. However, modern computing systems natively employ a horizontal data layout that preserves the cache line abstraction, leverages spatial locality in row buffers, and enables high memory throughput. This fundamental mismatch forces existing PUD architectures to frequently perform data layout transformations between horizontal and vertical formats, incurring significant performance, energy, and system integration overheads. Our goal is to eliminate data transposition overheads in PUD systems at low cost. To this end, we propose PipeDRAM, a PUD architecture that eliminates the need for runtime data layout transformation, enabling PUD operations directly over horizontally laid-out data. PipeDRAM's key ideas are to (i) deterministically reorganize bits inside each memory request to enable a PUD-friendly data placement within a DRAM array in a horizontal data layout, and (ii) employ a pipeline-based execution model that overlaps bit-dependent and bit-independent in-DRAM operations to exploit bit-level parallelism across the memory array. We compare PipeDRAM to different computing platforms. PipeDRAM provides (i) 11.8x, 11.8x, and 80.4x higher performance and (ii) 25.4x, 3.0x, and 38.0x lower energy consumption than three state-of-the-art PUD systems. PipeDRAM incurs low area cost on top of a DRAM chip (1.86%) and CPU die (0.05%). To enable further research on PUD systems, we open-source PipeDRAM at https://github.com/CMU-SAFARI/PipeDRAM.

CommentsExtended version of MICRO 2026 paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑