可配置分层全归约(Allreduce)
Configurable and Hierarchical Allreduce
浏览论文内容
中文总结 AI 辅助
针对MPI_Allreduce中小消息场景的性能挑战,提出可配置分层Allreduce算法CHIARA,在Polaris等超算平台及并行k均值应用中实现显著加速。
中文摘要 AI 辅助
MPI_Allreduce是大规模科学计算和分布式机器学习中性能最关键的集合通信操作之一,但中小消息场景仍具挑战性:延迟、同步深度以及域内与域间通信间的强硬件层级均会增加单次调用的开销。我们提出CHIARA,一种可配置分层Allreduce,它通过逻辑批-通道拓扑编码硬件层级,并执行分阶段调度,每次仅激活规约向量的有限部分。批间通信通过旋转根通道原语在多个进程间分布,避免集中式领导者。该工具还通过在Reduce-Scatter/Allgather边界保留通道对齐的中间布局,实现半组合式Rabenseifner风格Allreduce,消除冗余的域内重组。我们在Polaris、Aurora和Fugaku上评估该工具,相较于供应商MPI_Allreduce分别实现最高1.94倍、13.43倍、13.48倍的加速,在并行k均值应用中实现最高2.2倍的端到端加速。
英文摘要
MPI_Allreduce is among the most performance-critical collectives in large-scale scientific computing and distributed machine learning, yet the small- and medium-message regime remains challenging: latency, synchronization depth, and strong hardware hierarchy between intra- and inter-domain communication all compound per-invocation cost. We present CHIARA, a configurable hierarchical Allreduce that encodes hardware hierarchy through a logical batch-lane topology and executes a staged schedule in which only a bounded portion of the reduction vector is active at a time. Inter-batch communication is distributed across multiple ranks via a rotating-root lane primitive, avoiding centralized leaders. Tool further enables a semi-composed Rabenseifner-style Allreduce by preserving a lane-aligned intermediate layout across the Reduce-Scatter/Allgather boundary, eliminating redundant intra-domain reorganization. We evaluate Tool on Polaris, Aurora, and Fugaku, achieving speedups of up to 1.94x, 13.43x, and 13.48x over vendor MPI_Allreduce, and up to 2.2x end-to-end speedup in a parallel k-means application.