arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将半环动态规划编译为层级感知的3D-DRAM存内计算

Compiling Semi-Ring Dynamic Programming to Tier-Aware 3D-DRAM Processing-in-Memory

Mahbod Afarin, Tsung-Han Lu, Tajana Rosing

arXiv 2610.09156首次发表:更新:

发表机构

University of California San Diego(加州大学圣地亚哥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对3D-DRAM存内计算,提出GenMLIR编译器,将半环动态规划自动映射为层级感知的PIM代码,相比无编译支持和标准仿射编译器分别快达5.8/16.7倍和1.2-1.4倍,且代码量大幅减少。

AI 中文摘要

单片3D(M3D)DRAM上的存内计算(PIM)是应对数据密集型动态规划(DP)内存墙的有前景方案,但当前要发挥其性能需要手写内核:程序员必须选择分块大小,在非均匀延迟的内存层级间放置数据,在异构处理单元之间划分工作,并手动插入正确的广播和队列。我们提出GenMLIR,一个MLIR编译器,可自动实现面向层级感知3D PIM的这种映射。GenMLIR将半环广义网格更新(所有点对最短路径(APSP)和序列比对共享的代数形式)编码为一等IR抽象,并通过四个GenDRAM感知的遍组进行降级,这些遍组推导出分块、层级感知放置、块到PU的分配以及显式通信。我们还精确刻画了该抽象支持和不支持的DP递归关系。在GenDRAM架构上,GenMLIR生成的代码比无编译器支持的降级快最多5.8倍(APSP)和16.7倍(比对),比我们同时实现的、采用标准仿射分块的PIM编译器快1.2-1.4倍,达到可实现性能上限的90-100%。它用4-9行代码表达这些工作负载,而手写需要300-500行,且编译时间仅占运行时间的极小部分。

英文摘要

Processing-in-Memory (PIM) on monolithic 3D (M3D) DRAM is a promising answer to the memory wall for data-intensive dynamic programming (DP), yet extracting its performance today demands hand-written kernels: the programmer must pick a tile size, place data across non-uniform-latency memory tiers, partition work between heterogeneous processing units, and insert the right broadcasts and queues by hand. We present GenMLIR, an MLIR compiler that automates this mapping for tier-aware 3D PIM. GenMLIR encodes the semi-ring generalized grid update, the algebraic form shared by all-pairs shortest path (APSP) and sequence alignment, as a first-class IR abstraction, and lowers it through four GenDRAM-aware pass groups that derive blocked tiling, tier-aware placement, tile-to-PU assignment, and explicit communication. We also characterize precisely which DP recurrences the abstraction admits and which it does not. On the GenDRAM architecture, GenMLIR-generated code runs up to 5.8 (APSP) and 16.7 (alignment) faster than a lowering without compiler support, and 1.2-1.4 faster than a standard affine-tiling PIM compiler we also implement, reaching 90-100% of an achievable-performance bound. It expresses these workloads in 4-9 lines instead of 300-500 and compiles in a negligible fraction of runtime.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑