arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STORM:用于分布式内存粒子模拟的基于RDMA的蒙特卡罗传输方案

STORM: RDMA-based Monte Carlo Transport Scheme for Distributed-Memory Particle Simulations

Maor Mizrachi, Barak Raveh, Elad Steinberg

arXiv 2607.20639首次发表:更新:

AI 中文总结

该研究针对蒙特卡罗粒子传输在非结构化网格上通信限制可扩展性的问题,提出STORM开源库,用RDMA取代MPI语义,构建无锁通信层。实验表明其在多核心测试中效率高且有加速比,消除了蒙特卡罗传输缩放障碍,实现耦合辐射流体动力学。

AI 中文摘要

蒙特卡罗粒子传输能够通过处理多维几何结构、频率依赖性和移动介质而无需角度离散化,实现从核心坍缩超新星、中子星合并到吸积流等的高保真天体物理辐射和中微子模拟。然而,在非结构化网格上,跨等级通信限制了可扩展性。我们提出了STORM(通过单边远程内存实现可扩展传输),它是一个用于在通用网格、物理和边界条件下进行蒙特卡罗传输的开源库。STORM提供了一个无锁且与网格无关的通信层,用远程直接内存访问(RDMA)取代了MPI的匹配发送/接收语义。在对抗性均匀发射基准测试中,RDMA后端在高达13440个核心(每个网络适配器112个核心)时保持大于97%的弱缩放和大于88%的强缩放效率,比双边替代方案快1.14至1.27倍。在4480个等级的黑腔IMC基准测试中,它快1.41倍,因为MPI进度开销减少了6.1倍。通过将通信与物理模型和网格表示解耦,STORM消除了天体物理多物理代码中蒙特卡罗传输缩放的障碍,实现了在动态演化网格上的能量和角度分辨光子或中微子传输的耦合辐射流体动力学。

英文摘要

Monte Carlo particle transport enables high-fidelity astrophysical radiation and neutrino simulations - from core-collapse supernovae and neutron-star mergers to accretion flows - by handling multidimensional geometries, frequency dependence, and moving media without angular discretization. However, inter-rank communication limits scalability on unstructured meshes: standard two-sided MPI requires receivers to post receives and poll completions, creating per-iteration progress overhead that grows with the number of communication partners. Such problems have not demonstrated high scaling efficiency at $O(10^4)$ cores. We present STORM (Scalable Transport via One-sided Remote Memory), an open-source library for Monte Carlo transport on general meshes, physics, and boundary conditions. STORM provides a lock-free, mesh-independent communication layer that replaces MPI's matched-send/receive semantics with Remote Direct Memory Access (RDMA) - one-sided operations that write directly into a remote rank's memory without involving its CPU. Each rank pair shares a single-producer, single-consumer ring buffer; RDMA writes transfer particles while receivers remain passive. A two-sided MPI backend provides a portable fallback. In an adversarial uniform-emission benchmark, the RDMA backend sustains $>97\%$ weak-scaling and $>88\%$ strong-scaling efficiency up to 13,440~cores (112~cores per network adapter), with $1.14$-$1.27\times$ speedups over the two-sided alternative. In a Hohlraum IMC benchmark at 4480 ranks, it is $1.41\times$ faster because MPI progress overhead is reduced by $6.1\times$. By decoupling communication from physics models and mesh representations, STORM removes a barrier to scaling Monte Carlo transport in astrophysical multiphysics codes, enabling coupled radiation-hydrodynamics with energy- and angle-resolved photon or neutrino transport on dynamically evolving meshes at scale.

CommentsHas been submitted to ApJS. Comments are welcome: maormiz@cs.huji.ac.il

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑