AI 中文总结
VIPER是面向存内计算设计空间探索的架构感知性能评估框架,仅需一次主机分析,可快速准确估计性能,相比周期精确仿真大幅缩短时间,能揭示设备级评估遗漏的关键性能权衡。
AI 中文摘要
存内计算(Processing-in-Memory, PIM)通过在内存内部或附近执行计算,有望减少数据移动开销,但应用实际获得的加速效果高度依赖设计。不可卸载的主机执行、主机与PIM间的数据传输、有限的PIM容量以及设备编程延迟都会限制端到端加速比,因此快速的早期设计空间探索(Design-Space Exploration, DSE)至关重要。然而,现有的PIM评估方法存在局限:电路级和设备级工具无法捕捉这些端到端PIM性能因素,而周期精确的仿真对于迭代式DSE来说速度过慢。为解决这一问题,我们提出VIPER,这是一个统一、轻量且架构感知的PIM DSE性能评估框架。VIPER仅对主机执行进行一次分析,将测得的主机行为与感知PIM的分析引擎相结合,该引擎可在候选设计中遍历PIM端参数。它支持近内存计算(Processing Near Memory, PNM)和内存中计算(Processing Using Memory, PUM)两种模式,涵盖任务卸载和数据触发执行,通过捕捉主机-PIM传输、数组访问、内存内计算、设备编程延迟以及容量导致的分区,无需重复进行周期精确仿真,即可为迭代式DSE提供快速的架构感知性能估计。我们将VIPER与商用UPMEM系统及400余种周期精确的gem5配置进行验证。VIPER可预先预测UPMEM的卸载决策和盈亏平衡点,结合优化的传输模型,在DPU遍历中以12%的平均加速比误差捕捉测得的峰值及回落行为(256个DPU峰值时误差低至6%);对比gem5,VIPER实现了低于10%的误差,同时将评估时间从数小时缩短至一分钟以内。对UPMEM、ReRAM/FeFET交叉阵列及IMCRYPTO的案例研究表明,架构感知的DSE能够揭示设备级评估遗漏的关键性能权衡。
英文摘要
Processing-in-Memory (PIM) promises to reduce data movement overhead by executing computation in or near memory, but its realized application speedup remains highly design-dependent. Non-offloadable host execution, host-PIM transfers, limited PIM capacity, and device programming latency can limit end-to-end speedup, making fast early-stage design-space exploration (DSE) essential. However, existing PIM evaluation methods remain limited: circuit- and device-level tools cannot capture these end-to-end PIM performance factors, while cycle-accurate simulation is too slow for iterative DSE. To address this gap, we present VIPER, a unified, lightweight, and architecture-aware performance evaluation framework for PIM DSE. VIPER profiles host execution once and combines the measured host behavior with a PIM-aware analytical engine that sweeps PIM-side parameters across candidate designs. It supports both Processing Near Memory (PNM) and Processing Using Memory (PUM) under task-offloading and data-triggered execution by capturing host-PIM transfer, array access, in-memory computation, device programming latency, and capacity-induced partitioning, providing rapid architecture-aware performance estimates for iterative DSE without repeated cycle-accurate simulation. We validate VIPER against a commercial UPMEM system and more than 400 cycle-accurate gem5 configurations. VIPER predicts the UPMEM offloading decision and break-even region a priori, and, with a refined transfer model, captures the measured peak-and-rolloff behavior with 12\% mean speedup error across the DPU sweep (6\% up to the 256-DPU peak). Against gem5, VIPER achieves less than 10\% error while reducing evaluation time from hours to under one minute. Case studies of UPMEM, ReRAM/FeFET crossbars, and IMCRYPTO show that architecture-aware DSE reveals key performance trade-offs that device-level evaluation misses.
Comments14 pages, 12 figures. Source code available at https://github.com/Notre-Dame-HW-SW-Codesign-Lab/VIPER