异构内存架构的内存剖析与迁移
Memory Profiling and Migration for Heterogeneous Memory Architectures
浏览论文内容
中文总结 AI 辅助
提出SHAMBLES内核集成框架,通过低开销剖析和策略驱动迁移,在无需修改应用的情况下将热数据集中于HBM,在HPCG、DGEMM和Himeno上分别维持高达93.75%、99%和优于固定放置的性能,实现分层内存性能可移植性。
中文摘要 AI 辅助
结合高带宽内存(HBM)与普通DRAM的异构内存系统可以加速带宽受限的HPC工作负载,但当前的页面放置主要依赖手动调优或操作系统启发式方法,这些方法并非为多层动态而设计。我们提出了SHAMBLES,一个内核集成的框架,以低开销剖析应用程序的内存行为,并在无需修改应用程序的情况下跨层迁移数据。SHAMBLES提供了一个策略无关的接口和一个轻量级的用户空间运行时,支持可插拔策略(例如基于最近使用和频率的策略)以及用于受控研究的静态放置。日志模式提供分配和迁移的可重现时间线,以辅助分析。我们在一个将HBM和DDR暴露为NUMA节点的商用Linux系统上实现了SHAMBLES,并使用HPCG、DGEMM基准测试和Himeno模板迷你应用进行了评估。我们的设计和方法论展示了透明的、策略驱动的迁移如何响应变化的访问局部性,并在无需开发者干预的情况下将热数据集中在HBM中,为分层内存上的性能可移植性提供了一条实用路径。HPCG的结果表明,我们可以在仅将问题大小的40%保留在HBM中的情况下,维持高达全HBM基线性能的93.75%。DGEMM实验表明,SHAMBLES中的动态策略在仅将DGEMM矩阵占用空间的三分之一保留在HBM中的情况下,可持续维持高达全HBM性能的99%。对于Himeno,SHAMBLES表明快速层选择必须既感知工作负载又感知大小:在50%快速层预算下,对于L尺寸,它可以优于固定的全HBM和全DDR放置,而XL尺寸则回归到更偏向HBM。
英文摘要
Heterogeneous memory systems that combine high-bandwidth memory (HBM) with commodity DRAM can accelerate bandwidth-bound HPC workloads, but current page placement largely depends on manual tuning or OS heuristics not designed for multi-tier dynamics. We present SHAMBLES, a kernel-integrated framework that profiles application memory behavior at low overhead and migrates data across tiers without requiring application changes. SHAMBLES exposes a policy-agnostic interface and a lightweight user-space runtime with pluggable policies (e.g. recency and frequency based) as well as static placement for controlled studies. A logging mode provides reproducible timelines of allocations and migrations to aid analysis. We implement SHAMBLES on a commodity Linux system with HBM and DDR exposed as NUMA nodes and evaluate it with the HPCG, DGEMM benchmarks and Himeno stencil mini-app. Our design and methodology show how transparent, policy-driven migration can respond to changing access locality and concentrate hot data in HBM without developer intervention, offering a practical path to performance portability on tiered memory. Results from HPCG show that we can maintain up to 93.75% of the all-in-HBM baseline performance, while keeping only 40% of the problem size in the HBM. DGEMM experiments show that dynamic policies in SHAMBLES sustain up to 99% of the all-in-HBM performance, while keeping only one third of the DGEMM matrix footprint in HBM. For Himeno, SHAMBLES shows that fast-tier selection must be both workload-aware and size-aware: with a 50% fast-tier budget, it can outperform fixed all-in-HBM and all-in-DDR placements for the L size, while the XL size shifts back toward HBM.
发表机构
- Institute of Computer Science Foundation for Research and Technology - Hellas (FORTH)(希腊基础研究基金会计算机科学研究所)
机构由 AI 辅助整理,请以论文原文为准。