arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Tiga:大规模图消息传递的编译

Tiga: Compiling Graph Message Passing at Scale

Mingyuan Chi

arXiv 2609.24802首次发表:更新:

AI 中文总结

Tiga是一个即时编译器,通过多级中间表示和流式分区执行,在内存受限设备上高效编译图消息传递,减少运行时和内存使用,支持分布式执行。

AI 中文摘要

图消息传递为表达学习算法、物理模拟和数值求解器提供了一种常见方式。其高效执行依赖于交互结构和数据移动,而当程序以张量操作序列表达时,这些因素可能被掩盖。在内存受限的系统(如笔记本电脑)上,物化连接性和中间消息也可能耗尽设备内存。我们提出了Tiga,一个即时编译器,它将消息传递程序的定义与其交互的遍历、计算和存储方式分离开来。Tiga在多级中间表示中保留图关系和归约器代数,从而实现遍历特化、关系生成与聚合的融合,以及反向模式自动微分。后端特定的降级面向CPU和GPU。其运行时通过从磁盘流式传输图分区(通过主机内存,使用页大小的设备暂存缓冲区)将执行扩展到设备内存容量之外;分区所有权和halo交换将相同的编程模型扩展到分布式执行。Python接口与普通PyTorch张量和autograd互操作,用于前向和反向计算。数值检查验证了可微分工作负载的输出和梯度。一个可微分的几何工作负载表征了融合前向和反向执行的内存-时间权衡。与匹配的Torch和PyTorch Geometric基线的评估表明,在生成关系工作负载上减少了运行时和设备内存使用,而卸载的前向执行在单个内存受限的GPU上处理了十亿边图。异构设备上的测量进一步表征了分布式执行的通信和负载均衡成本。

英文摘要

Graph message passing offers a common way to express learning algorithms, physical simulations, and numerical solvers. Efficient execution depends on interaction structure and data movement, which can be obscured when a program is expressed as a sequence of tensor operations. On memory-constrained systems such as laptops, materializing connectivity and intermediate messages can also exhaust device memory. We present Tiga, a just-in-time compiler that separates the definition of a message-passing program from how its interactions are traversed, computed, and stored. Tiga preserves graph relations and reducer algebra in a multi-level intermediate representation, enabling traversal specialization, fusion of relation generation with aggregation, and reverse-mode automatic differentiation. Backend-specific lowering targets CPUs and GPUs. Its runtime extends execution beyond device-memory capacity by streaming graph partitions from disk through host memory with page-sized device staging buffers; partition ownership and halo exchange extend the same programming model to distributed execution. A Python interface interoperates with ordinary PyTorch tensors and autograd for forward and backward computation. Numerical checks validate outputs and gradients for differentiable workloads. A differentiated geometric workload characterizes the memory--time tradeoff of fused forward and backward execution. Evaluation against matched Torch and PyTorch Geometric baselines demonstrates reduced runtime and device-memory use for generated-relation workloads, while offloaded forward execution processes billion-edge graphs on a single memory-limited GPU. Measurements on heterogeneous devices further characterize the communication and load-balance costs of distributed execution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑