发表机构
Yonsei University; Kiel University(延世大学; 基尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Torch-PIM是一个编译器框架,利用配置文件引导优化在循环嵌套级别自动决定主机与PIM的卸载,在多种PIM配置下对张量操作符、MLP、Attention及大模型实现高达8.6倍的加速。
AI 中文摘要
现代深度学习(DL)工作负载受限于数据移动,而存内计算(PIM)通过将计算单元放置在内存附近来针对这一瓶颈。然而,PyTorch和其他DL框架缺乏编译器支持来对其下层的代码做出此决策:现有的卸载框架面向手写的C/C++程序,而那些针对DL的框架在降低之前将候选集固定为操作符类型列表。我们提出了Torch-PIM,一个编译器框架,它使用配置文件引导优化(PGO)来决定在渐进式降低所物化的循环嵌套中主机与PIM的放置。流水线发出的每个并行循环嵌套都进入候选空间,并且每个循环嵌套在两个阶段进行评估:其承载的工作量以及其内存受限性。评估所消耗的每个数量都在主机上进行剖析或从代码的多级中间表示(MLIR)中获取。在32到128核的PIM配置中,Torch-PIM的卸载决策在张量操作符上实现了高达8.6倍、MLP上2.9倍、Attention上4.4倍、GPT-J-6B上5.1倍以及LLaMA-7B上3.6倍的加速比,相较于仅CPU执行。
英文摘要
Modern deep learning (DL) workloads are limited by data movement, and processing-in-memory (PIM) targets this bottleneck by placing compute units near the memory. However, PyTorch and other DL frameworks lack compiler support for making this decision on the code they lower: existing offloading frameworks target hand-written C/C++ programs, while those that address DL fix the candidate set to a list of operator types before lowering. We present Torch-PIM, a compiler framework that uses profile-guided optimization (PGO) to decide host-versus-PIM placement over the loop nests that progressive lowering materializes. Every parallel loop nest the pipeline emits enters the candidate space, and each is assessed in two stages: the amount of work it carries, and its memory boundedness. Every quantity the assessment consumes is profiled on the host or obtained from the multi-level intermediate representation (MLIR) of the code. Across PIM configurations of 32 to 128 cores, Torch-PIM's offloading decisions yield speedups of up to 8.6x on tensor operators, 2.9x on MLP, 4.4x on Attention, 5.1x on GPT-J-6B, and 3.6x on LLaMA-7B over CPU-only execution.
Comments12 pages, 3 figures, 3 tables