arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29219cs.PL

调度是可解符号:数据流架构上瓦片程序的免调优编译

Schedules Are Solvable Symbols: Tuning-Free Compilation of Tile Programs on Dataflow Architectures

Heru Wang, Wei Li, Zhenyu Bai, Tulika Mitra

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Loom,一种免调优符号编译框架,通过将瓦片程序编译视为硬件显式静态优化问题,在数据流架构上联合求解调度与块大小,匹配或超越供应商优化库。

中文摘要 AI 辅助

现代AI和高性能计算加速器日益暴露数据流特性:软件可见的数据移动和重叠机制,例如通过片上网络的核间通信和核内异步流水线。这些特性将调度责任从硬件转移到编译器,并且由于放置、移动和同步变得软件可见,它们也使静态调度的性能变得可预测。然而,在此类硬件上实现高性能仍依赖于供应商精心设计的核库或基于性能剖析的自动调优,其嵌入的专家知识在不同架构和算法间迁移效果不佳。我们提出Loom,一个面向空间数据流架构上基于瓦片的SPMD程序的免调优符号编译框架。核心思想是将基于瓦片的SPMD编译视为一个硬件显式的静态优化问题。Loom枚举离散的空间映射和通信候选,同时保持值参数(如瓦片因子和流水线旋钮)在每个候选中为符号形式。从显式硬件描述中,它推导出符号合法性约束和延迟表达式,为每个调度候选制定一个CP-SAT问题,并在编译时联合求解核间数据流、核内异步调度和块大小。在两代Tenstorrent硬件(Wormhole和Blackhole)上,Loom在GEMM、Flash Attention和Flash Decode上匹配或超过供应商优化的TTNN库,开箱即用,无需针对形状的剖析或基于剖析的平台特定调度调优。这些结果表明,硬件派生的符号编译为空间数据流架构提供了一种可重定向的替代基于剖析调优的方法,同时通过将优化决策追溯到源级符号保持可解释性。

英文摘要

Modern AI and HPC accelerators increasingly expose dataflow features: software-visible mechanisms for data movement and overlap, such as inter-core communication through the on-chip network and intra-core asynchronous pipelining. These features shift scheduling responsibility from hardware to the compiler, and because placement, movement, and synchronization become software-visible, they also make the performance of static schedules predictable. Yet high performance on such hardware still relies on vendor-engineered kernel libraries or profile-based auto-tuning, whose embedded expert knowledge transfers poorly across architectures and algorithms. We present Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures. The central idea is to treat tile-based SPMD compilation as a hardware-explicit static optimization problem. Loom enumerates discrete spatial-mapping and communication candidates while keeping value parameters, such as tiling factors and pipeline knobs, symbolic within each candidate. From an explicit hardware description, it derives symbolic legality constraints and latency expressions, formulates one CP-SAT problem per schedule candidate, and jointly solves inter-core dataflow, intra-core asynchronous scheduling, and block sizes at compile time. On two Tenstorrent generations, Wormhole and Blackhole, Loom matches or exceeds the vendor-optimized TTNN library on GEMM, Flash Attention, and Flash Decode, out of the box and without per-shape profiling or profile-based platform-specific schedule tuning. These results suggest that hardware-derived symbolic compilation provides a retargetable alternative to profiling-based tuning for spatial dataflow architectures while remaining interpretable by keeping optimization decisions traceable to source-level symbols.

发表机构

  • School of Computing, National University of Singapore(新加坡国立大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

↑