arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PIMID:一种用于内存处理的具有复杂性和多样性的全系统模拟器

PIMID: A Full-System Simulator with Intricacy and Diversity for Processing-in-Memory

Yuan He, Masaaki Kondo, Galen M. Shipman, Jered B. Dominguez-Trujillo, Shigeki Tomishima, Kazi Asifuzzaman

arXiv 2607.24196首次发表:更新:

发表机构

RIKEN Center for Computational Science; Keio University; Los Alamos National Laboratory; Oak Ridge National Laboratory(理化学研究所计算科学中心; 庆应义塾大学; 洛斯阿拉莫斯国家实验室; 橡树岭国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究内存处理设计空间,提出全系统模拟器PIMID,支持共享内存和消息传递模型,能在多种内存技术上运行,可放置处理元素、扫描相关参数并定价内存网络,通过协同模拟给出时间和能量分解,揭示多种内存处理特性及PIMID插件接口优势。

AI 中文摘要

内存处理通过将计算与内存共置来解决内存墙问题,但由于实际的内存处理硬件仍然稀缺,模拟是探索内存处理设计空间的主要方式。现有的内存处理模拟器只覆盖了该空间的一部分:通常模拟单一内存技术,将处理元素固定在内存层次结构的一个级别,支持单一执行模型,并止于设备边界。因此,我们提出了PIMID,一个执行和跟踪驱动的全系统模拟器,它在一个工具中弥补了这些差距。PIMID支持共享内存和消息传递执行模型,可在11种内存技术(7种DRAM标准、SRAM和3种非易失性存储器)上并行运行OpenMP和MPI注释的并行代码;它可以将处理元素从子阵列放置到逻辑芯片,扫描处理元素数量和核心模型保真度,并根据测量的拥塞情况对每种技术的内存网络进行定价。其单进程主机-设备协同模拟解决了仅设备工具无法产生的端到端时间和能量分解(主机准备、设备计算和显式边界电荷)。在由此产生的双执行模型数据集中,PIMID表明,仅内存技术就能使执行时间变化超过一个数量级,并且最佳的主机主内存并非最佳的内存处理基板;规则内核随着处理元素数量超线性扩展,因为内存带宽与计算共同扩展;消息传递下的图遍历在共享内存下不存在集体通信墙;并且在全系统范围内,卸载仅在带宽级内存上以时间换能量:共享内存卸载在HBM3上节省能量,而16核主机在每个端到端时间上都占优。PIMID的插件接口允许随着内存处理技术的发展,通过标准化的YAML规范添加新的引擎和模型。

英文摘要

Processing-in-Memory addresses the memory wall by co-locating computation with memory, but because real PIM hardware remains scarce, simulation is the primary way to explore the PIM design space. Yet existing PIM simulators each cover only part of that space: they typically model a single memory technology, fix processing elements at one level of the memory hierarchy, support a single execution model, and stop at the device boundary. We therefore present PIMID, an execution- and trace-driven full-system simulator that closes these gaps in one tool. PIMID supports both the shared-memory and message-passing execution models, running annotated parallel code in OpenMP and MPI side by side across eleven memory technologies (seven DRAM standards, SRAM, and three non-volatile memories); it places PEs anywhere from subarrays to logic dies, sweeps PE count and core-model fidelity, and prices the in-memory network per technology from measured congestion. Its single-process host-device co-simulation resolves an end-to-end time and energy breakdown (host preparation, device compute, and explicit boundary charges) that device-only tools cannot produce. Across the resulting dual-execution-model dataset, PIMID shows that the memory technology alone moves execution time by more than an order of magnitude and that the best host main memory is not the best PIM substrate; that regular kernels scale superlinearly with PE count as in-memory bandwidth co-scales with compute; that graph traversal under message-passing hits a collective-communication wall absent under shared memory; and that at full-system scope the offload trades time for energy only on the bandwidth-class memory: shared-memory offload saves energy on HBM3 while a 16-core host keeps every end-to-end time win. PIMID's plugin interfaces let new engines and models be added through standardized YAML specifications as PIM technology evolves.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑