发表机构
Syracuse University; Friedrich-Alexander-Universität Erlangen-Nürnberg; TU Dresden(雪城大学; 埃尔朗根-纽伦堡大学; 德累斯顿工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RAPID通过在DRAM子阵列中引入迁移单元和反相单元,实现行并行、位并行算术处理,消除数据布局转换,在19个MLPerf基准上比SIMDRAM提升5.9倍端到端性能。
AI 中文摘要
使用内存处理(PUM)架构直接在DRAM内执行计算,以减少内存与处理器之间昂贵的数据移动。由于电荷共享操作局限于单个位线,现有的DRAM-PUM架构将数据重组为列导向、位串行的表示形式。这种组织方式与常规处理器和加速器使用的行导向、字并行布局根本不相容,导致计算在PUM和常规执行之间切换时需要进行昂贵的数据布局转换。在本文中,我们提出了RAPID,一种行并行算术处理内存(Processing-In-DRAM)架构。RAPID通过两种轻量级扩展增强了DRAM子阵列:迁移单元,实现相邻位线之间的局部水平数据移动;反相单元,提供高效的阵列内逻辑反相。这些原语使RAPID能够直接操作行并行、位并行的数据,在保持CPU兼容布局的同时,利用DRAM子阵列的大规模并行性。特别是,RAPID证明了局部水平通信足以实现浅层算术网络和高效的乘法并行归约,从而降低算术延迟,同时保持吞吐量并消除昂贵的数据布局转换,同时保持常规DRAM阵列组织。我们通过详细的晶体管级布局和SPICE验证的电路仿真,证明了用迁移和反相单元增强DRAM子阵列的可行性和开销。使用RAPID编译器,可以评估性能和数据重组权衡,以确保在CPU和PUM组合执行中获得最佳性能。在19个MLPerf基准测试上评估RAPID,与SIMDRAM相比,DDR4 PUM执行实现了5.9倍的端到端性能提升。
英文摘要
Processing-using-memory (PUM) architectures perform computation directly within DRAM to reduce costly data movement between memory and processors. Because charge-sharing operations are confined to individual bitlines, existing DRAM-PUM architectures reorganize data into column-oriented, bit-serial representations. This organization is fundamentally incompatible with the row-oriented, word-parallel layouts used by conventional processors and accelerators, requiring expensive data-layout transformations whenever computation transitions between PUM and conventional execution. In this paper, we present RAPID, a Row-parallel Arithmetic Processing-In-DRAM architecture. RAPID augments the DRAM subarray with two lightweight extensions: migration cells that enable localized horizontal data movement between neighboring bitlines and inversion cells that provide efficient in-array logical inversion. These primitives enable RAPID to operate directly on row-parallel, bit-parallel data, preserving CPU-compatible layouts while exploiting the massive parallelism of the DRAM subarray. In particular, RAPID demonstrates that localized horizontal communication is sufficient to realize shallow arithmetic networks and efficient parallel reduction for multiplication, reducing arithmetic latency while preserving throughput and eliminating costly data-layout transformations, all while maintaining the conventional DRAM array organization. We demonstrate the feasibility and overhead of augmenting DRAM subarrays with migration and inversion cells through detailed transistor-level layout and SPICE-validated circuit simulations. Using the RAPID compiler it is possible to evaluate the performance and data reorganization tradeoffs to ensure the best execution across combined CPU and PUM. Evaluating RAPID on 19 MLPerf benchmarks, there is a 5.9x higher end-to-end performance compared to SIMDRAM for DDR4 PUM execution.