arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用MPI应用中计算密集与内存密集内核的相互作用

Exploiting the Interplay of Compute- and Memory-Bound kernels in MPI Applications

Ayesha Afzal, Krishna Manda, Georg Hager

arXiv 2610.01587首次发表:更新:

发表机构

University of Bonn(波恩大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现MPI应用中通信停顿可缓解内存带宽争用,通过并行光流求解器展示去同步化带来的加速,并开发微基准测试与模拟器验证,表明自然去同步化是架构感知优化。

AI 中文摘要

并行应用通常设计为同步、锁步执行,将通信停顿视为性能隐患。然而,在通信量少且无频繁同步点、并在计算密集与内存密集执行之间交替的应用中,MPI通信停顿可无意中缓解内存带宽争用。我们使用并行光流求解器演示了这一点,该求解器结合了计算密集的光线追踪内核与内存密集的光流求解器内核,且进程间通信可忽略不计。该程序通过去同步化以及计算密集与内存密集阶段间的自动重叠展现出显著加速,表明自然去同步化是一种架构感知的优化。当在ccNUMA域上并发执行内存密集阶段的进程数接近带宽饱和点时,可实现最优加速。我们还展示了一个案例,即使用MPI异步推进减少通信开销会显著降低性能,因为它允许过多秩同时争用内存带宽。为在更可控条件下研究动态,我们开发了一个可调的双内核微基准测试,并表明需要显著的应用程序或系统噪声(自然的或注入的)才能实现完全去同步化。最后,我们还使用一个带宽感知的基于模型的模拟器验证了这些结果。

英文摘要

Parallel applications are often designed for synchronous, lock-step execution, treating communication stalls as performance hazards. Yet, in a communication-light application without frequent synchronization points that alternates between compute-bound memory-bound execution, an MPI communication stall can act as an unintentional relief on memory-bandwidth contention. We demonstrate this using a Parallel Optical Flow Solver, which combines a compute-bound Ray Tracing kernel with a memory-bound Optical Flow Solver kernel and negligible inter-process communication. This program shows considerable speedup via desynchronization and automatic overlap between compute- and memory-bound phases, showing that natural desynchronization is an architecture-aware optimization. An optimal speedup is achieved when the number of processes concurrently executing the memory-bound phase on a ccNUMA domain is near the bandwidth saturation point. We also show a case where reducing communication overhead using MPI asynchronous progress significantly degrades performance because it allows too many ranks to contend for memory bandwidth simultaneously. In order to study the dynamics under more controlled conditions, we develop a tunable dual-kernel microbenchmark, with which we show that significant application or system noise (natural or injected) is required to achieve full desynchronization. Finally, we also validate these results using a bandwidth-aware, model-based simulator.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑