arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于RNN工作负载平衡的三维多细胞生长可扩展多GPU模拟

Scalable Multi-GPU Simulation of 3D Multicellular Growth with RNN-Based Workload Balancing

Matvey Moisseyev, Huijing Du, Dandan Zheng, Chi Zhang, Hongfeng Yu

arXiv 2608.25890首次发表:更新:

发表机构

School of Computing, University of Nebraska–Lincoln; Department of Mathematics, University of Nebraska–Lincoln; Department of Radiation Oncology, University of Rochester Medical Center; School of Biological Sciences, University of Nebraska–Lincoln; Holland Computing Center, University of Nebraska–Lincoln(内布拉斯加大学林肯分校计算学院; 内布拉斯加大学林肯分校数学系; 罗切斯特医学中心放射肿瘤学系; 内布拉斯加大学林肯分校生物科学学院; 内布拉斯加大学林肯分校霍兰德计算中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出结合GPU加速等技术的多GPU模拟框架,引入RNN负载平衡控制器,实现三维多细胞生长模拟的高效可扩展计算,提升负载平衡效果并降低运行时间。

AI 中文摘要

基于亚细胞元件模型(SEMs)的详细多细胞生长模拟可捕捉复杂的组织发育,但其元件级相互作用带来了巨大的计算成本。本研究提出了一种用于三维多细胞生长模拟的可扩展多GPU框架,该框架结合了GPU加速、空间分箱、域分解以及工作负载感知分区。细胞的运动、生长和分裂会持续重塑空间工作负载分布,导致初始平衡的分区随时间变得低效。为解决该问题,我们引入了一种基于RNN的负载平衡控制器,该控制器观察各进程的近期执行时间和分区状态,并学习对反应性边界调整规则的残差修正。该控制器在负载平衡循环的可微代理中通过随机工作负载动力学进行离线训练,无需 measured 执行轨迹即可训练。我们从单GPU加速、多GPU计算扩展、控制器级负载平衡行为以及端到端模拟性能等方面对该框架进行评估,与静态分区、反应性负载平衡以及传统时间序列预测基线进行了比较。一个代表性的胚胎表皮发育用例进一步展示了该框架所针对的时空演化工作负载类型。在我们的评估中,结合空间分箱的GPU加速相比串行CPU基线将相互作用计算加速了约三个数量级;RNN引导的负载平衡将平均全局不平衡从静态分区下的11.3%降至3.5%,相比静态分区降低了9.0%的端到端运行时间,与反应性基线相比减少了7.7倍的切片迁移,表明基于历史的控制可在避免不必要的重新分区的同时改善工作负载平衡。

英文摘要

Detailed multicellular growth simulations based on subcellular element models (SEMs) can capture complex tissue development, but their element-level interactions impose substantial computational cost. This work presents a scalable multi-GPU framework for 3D multicellular growth simulation that combines GPU acceleration, spatial binning, domain decomposition, and workload-aware partitioning. Cell movement, growth, and division continuously reshape the spatial workload distribution, causing initially balanced partitions to become inefficient over time. To address this, we introduce an RNN-based load-balancing controller that observes recent per-rank execution times and partition states and learns residual corrections to a reactive boundary-adjustment rule. The controller is trained offline in a differentiable surrogate of the load-balancing loop with randomized workload dynamics, requiring no measured execution traces for training. We evaluate the framework in terms of single-GPU acceleration, multi-GPU computation scaling, controller-level load-balancing behavior, and end-to-end simulation performance, with comparisons against static partitioning, reactive load balancing, and conventional time-series prediction baselines. A representative embryonic epidermal development use case further demonstrates the type of spatially and temporally evolving workload targeted by the framework. In our evaluation, GPU acceleration with spatial binning accelerates the interaction computation by roughly three orders of magnitude over a serial CPU baseline. RNN-guided load balancing reduces the mean global imbalance from 11.3% under static partitioning to 3.5%, lowers end-to-end runtime by 9.0% relative to static partitioning, and reduces slice migration by 7.7x compared with the reactive baseline, showing that history-aware control can improve workload balance while avoiding unnecessary repartitioning.

Comments10 pages, 9 figures. Currently under review for IEEE BigData 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑