arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28724cs.PF

规范并行形式:并行化编译器与智能体优化器的基底

The Canonical Parallel Form as a Substrate for Parallelizing Compilers and Agentic Optimizers

Yakup Koray Budanaz, Pratyai Mazumder, Alexandru Calotoiu, Torsten Hoefler

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出规范并行形式(CPF),通过移除不必要的排序约束并分层推导并行性,在AMD MI300A上显著加速循环内核,同时降低智能体优化器的令牌成本。

中文摘要 AI 辅助

命令式代码固定了计算并不需要的执行顺序,而并行化编译器必须证明该顺序的哪些部分可以被移除。我们引入了规范并行形式(CPF),这是一种与设备无关的程序状态,其中我们分析证明不必要的所有排序约束均已被移除。CPF通过保持输出的规范化、提升张量收缩等语义操作,以及按可判定性排序的三个层级推导并行性来获得:语法下标测试、整数集上的精确仿射依赖测试,以及非线性整数算术上的SMT查询,后者还允许在运行时检查之后进行并行。启发式方法随后可为每种架构定制规范形式。在AMD MI300A上的248个循环级推理内核中,CPF在其24个Zen 4核心上达到Numba的4.4倍速度,在其CDNA 3 GPU上达到25.4倍。与其他自动并行化优化器相比,CPF在CPU上比DaCe自身的自动并行化器快2.9倍,在GPU上快8.7倍,比Pluto快1.9倍,在内核的仿射子集上比PPCG快1.5倍。由于流水线是确定性的,CPF还可作为智能体的起始源代码,将每个内核的令牌成本降低最多2.72倍,同时保持所实现的加速不变,因为编码智能体对并行性的推理更少。

英文摘要

Imperative code fixes an execution order the computation does not require, and a parallelizing compiler must prove which parts of that order it can remove. We introduce the Canonical Parallel Form (CPF), a device-neutral program state from which every ordering constraint our analyses prove unnecessary has been removed. CPF is reached by output-preserving normalization, by lifting semantic operations such as tensor contractions, and by deriving parallelism in three levels ordered by decidability: syntactic subscript tests, exact affine dependence tests over integer sets, and SMT queries over non-linear integer arithmetic, which also admit parallelism guarded behind a runtime check. Heuristics can then specialize the canonical form for each architecture. Across 248 loop-level reasoning kernels on an AMD MI300A, CPF reaches 4.4x over Numba on its 24 Zen 4 cores and 25.4x on its CDNA 3 GPU. Against the other auto-parallelizing optimizers, CPF is 2.9x faster on the CPU and 8.7x on the GPU than DaCe's own auto-parallelizer, 1.9x faster than Pluto, and 1.5x faster than PPCG on the affine subset of the kernels. Because the pipeline is deterministic, CPF also serves as an agent's starting source, cutting the token cost per kernel by up to a factor of 2.72x while leaving the achieved speed-up unchanged, since the coding agents reason less about parallelism.

发表机构

  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑