arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多引擎AI加速器的无障碍同步算法

A Barrier-Free Synchronization Algorithm for Multi-Engine AI Accelerators

Chungha Sung, Nikil V. Shyamsunder, Hanliang Zhang, Daniel Kroening, Joonwon Choi

arXiv 2608.13757首次发表:更新:

AI 中文总结

针对多引擎AI加速器的屏障同步缺陷,提出基于运行时动态阈值的无障碍同步算法,在ML内核和微基准上实现延迟降低、加速提升,且经Lean证明助手验证正确性。

AI 中文摘要

AWS Trainium等多引擎AI加速器包含并行执行的专用计算引擎,编译器必须同步它们之间的数据依赖关系。对于直线代码,这很简单:每个依赖关系简化为等待指令完成的阈值计数,该计数由编译器静态计算。循环无法采用此类静态阈值;简单的解决方案是在迭代边界插入全引擎屏障,重置同步状态以便将每个循环体视为直线代码,但会以牺牲并行性为代价。我们提出一种无障碍同步算法,该算法可通过任意嵌套、动态有界的结构化控制流精确强制每个依赖关系。关键思路是从跟踪的循环迭代次数在运行时计算动态阈值。我们在AWS Neuron ISA级别将其实现为编译器后端 passes。在一组ML内核上,与基于屏障的基线相比,它将延迟降低10%-45%,在同步受限的微基准上实现3.3倍加速,且通常与手动调优的手动分配相当或更优。过早发出消费者会违反其依赖关系,而过晚发出则会不必要地暂停执行。我们正式表征了正确性所需的最小同步,并在Lean证明助手中通过互模拟验证了我们的算法符合该标准。

英文摘要

Multi-engine AI accelerators such as AWS Trainium comprise specialized compute engines that execute in parallel, and the compiler must synchronize the data dependencies between them. For straight-line code this is simple: each dependency reduces to waiting for a threshold count of instruction completions, which the compiler computes statically. Loops admit no such static threshold; a simple solution inserts all-engine barriers at iteration boundaries, resetting synchronization state so each loop body can be treated as straight-line, at the cost of parallelism. We present a barrier-free synchronization algorithm that instead enforces each dependency precisely across structured control flow with arbitrarily nested, dynamically bounded loops. The key idea is to compute dynamic thresholds at runtime from tracked loop iteration counts. We implemented it as a compiler backend pass at the AWS Neuron ISA level. On a suite of ML kernels, it reduces latency 10-45% relative to the barrier-based baseline, achieves a 3.3x speedup on a synchronization-bound microbenchmark, and often matches or exceeds hand-tuned manual allocation. Issuing a consumer too early violates its dependency, while issuing too late unnecessarily stalls execution. We formally characterize the minimum synchronization required for correctness and verify in the Lean proof assistant, via bisimulation, that our algorithm meets this criterion.

CommentsTo appear in the 2027 IEEE/ACM International Symposium on Code Generation and Optimization (CGO '27). 16 pages, 8 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑