每个内核都是连接:Einsummable中AI计算的自动多GPU并行
Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in Einsummable
浏览论文内容
中文总结 AI 辅助
Einsummable是自动将AI计算分布到多GPU服务器的原型系统,通过连接-聚合规格选择分解方式,在8个GPU A100服务器上处理LLaMA Transformer块的性能优于手动调优的PyTorch和vLLM。
中文摘要 AI 辅助
将AI计算分布到多GPU服务器的各个GPU上是AI系统领域的核心问题之一。我们提出了Einsummable,一个原型系统,它接受类似PyTorch的AI计算描述,并自动将其分布到多GPU服务器上,无需程序员编写设备分配、分片标注或通信操作。Einsummable将每个操作建模为关系连接,随后对张量关系进行聚合,其中元组包含子张量。每个操作通过我们所谓的“连接-聚合规格(join-agg specs)”公开其可能的分解方式。随后,优化器会选择整个计算过程中的分解方式,以最小化通信成本代理值。由于它搜索的是分解方式而非一组命名策略,Einsummable能发现基于网格的自动并行器无法实现的执行计划。每个分解后的操作通过合成交换程序实现,该程序是Volcano交换算子的拓扑感知泛化。Einsummable不调用任何固定的集合通信操作:所有通信和聚合都是在编译时派生的专用操作。尽管完全自动化,Einsummable的性能却能超过定制设计的实现。例如,在8个GPU的A100服务器上处理LLaMA Transformer块时,Einsummable的几何平均运行时间为8.97毫秒,而手动调优的PyTorch为13.80毫秒,vLLM为15.90毫秒。
英文摘要
Distributing an AI computation across the GPUs of a multi-GPU server is one of the central problems in systems-for-AI. We present Einsummable, a prototype system that accepts a PyTorch-like description of an AI computation and automatically distributes it across a multi-GPU server, with no device assignments, sharding annotations, or communication operations written by the programmer. Einsummable models every operation as a relational join followed by an aggregation over tensor relations, in which the tuples contain sub-tensors. Each operation exposes its possible decompositions through what we call "join-agg specs". An optimizer then selects decompositions across the whole computation to minimize a communication-cost proxy. Because it searches decompositions rather than a menu of named strategies, Einsummable discovers plans that mesh-based auto-parallelizers cannot. Each decomposed operation is implemented by synthesizing an exchange program, which is a topology-aware generalization of Volcano's exchange operator. Einsummable invokes no canned collectives: all communication and aggregation is special-purpose, derived at compile time. Despite being fully automatic, Einsummable can outperform custom-designed implementations. For example, on LLaMA transformer blocks on an eight-GPU A100 server, Einsummable achieves a geometric-mean runtime of 8.97 ms, versus 13.80 ms for hand-tuned PyTorch and 15.90 ms for vLLM.
发表机构
- Rice University(莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。