arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于 gem5 的计算型 DRAM 仿真框架

A gem5-based Simulation Framework for Computing-in-DRAM

Alexander Kusnezoff, João Paulo C. de Lima, Jeronimo Castrillon, Asif Ali Khan

arXiv 2610.08186首次发表:更新:

发表机构

TU Dresden; Friedrich-Alexander-Universität Erlangen-Nürnberg; Paderborn University, Germany(德累斯顿工业大学; 埃尔朗根-纽伦堡大学; 帕德博恩大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对计算型 DRAM 缺乏全应用仿真工具的问题,提出基于最新 gem5 的 gem5-CIMD 框架,扩展指令集至十六种 CIM 操作,实现子阵列感知映射与刷新建模,并公开源码。

AI 中文摘要

使用 DRAM 的内存计算(CIMD)通过在 DRAM 子阵列内直接执行计算,为包含批量位操作的内存受限工作负载带来了显著的能效和吞吐量提升。然而,实现 CIMD 需要重新设计内存控制器,并仔细将操作数映射到内存阵列。目前,存在基于 FPGA 的精确测试平台,但它们成本高昂且劳动密集,而开源模拟器大多基于踪迹,无法执行完整的应用程序运行。唯一可用的全应用 CIMD 模拟器(随 MIMDRAM 提供)构建在过时的 gem5 版本及其工具链之上。我们提出了 gem5-CIMD,一个基于最新版 gem5 构建的全应用 CIMD 仿真框架。我们将模拟的指令集从四种位操作扩展到十六种 CIM 指令,涵盖算术、关系、条件和实用操作,支持多种数据类型和位宽(4 至 64 位)。一个完整的内存管理栈,包括多级大页支持、CIM 感知分配器和子阵列感知地址映射方案,确保 CIM 操作数位于同一 DRAM 子阵列中。此外,gem5-CIMD 感知 DRAM 刷新间隔,并在需要时忠实地建模和调度强制刷新操作。我们提供了一个 CIM 标准库作为编译器目标,并在多个案例研究中评估了 gem5-CIMD,包括一个端到端的 KNN 工作负载。模拟器源代码(包括评估的工作负载)已公开可用。

英文摘要

Computing-in-Memory using DRAM (CIMD) has demonstrated substantial energy and throughput gains for memory-bound workloads consisting of bulk-bitwise operations, by performing computation directly within DRAM subarrays. Realizing CIMD, however, requires a redesign of the memory controller and careful mapping of operands onto the memory arrays. Presently, accurate FPGA-based testbeds exist, but they are costly and labor-intensive, while open-source simulators are largely trace-based and cannot execute full application runs. The only full-application CIMD simulator available, included with MIMDRAM, is built on an outdated version of gem5 and its toolchains. We present gem5-CIMD, a full-application CIMD simulation framework built on the latest version of gem5. We extend the simulated instruction set from four bitwise operations to sixteen CIM instructions spanning arithmetic, relational, conditional, and utility operations across a range of data types and bitwidths (4 to 64 bits). A complete memory-management stack comprising multi-level huge-page support, a CIM-aware allocator, and a subarray-aware address-mapping scheme ensures that CIM operands are co-located in the same DRAM subarray. In addition, gem5-CIMD is aware of the DRAM refresh interval and faithfully models and schedules the mandatory refresh operations when needed. We provide a CIM standard library intended as a compiler target and evaluate gem5-CIMD on multiple case studies including an end-to-end KNN workload. The simulator sources, including the evaluated workloads, are publicly available.

Comments8 pages, 4 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑