arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02851cs.DC

ByteSplat:通过GPU内与GPU间通信减少实现高效分布式3D高斯溅射训练

ByteSplat: Efficient Distributed 3D Gaussian Splatting Training via Intra- and Inter-GPU communication reduction

Shuo Wu, He Zhu, Han Zhao, Xiaohui Zhang, Yaqian Zhao, Hui Wei, Ruyang Li, Hongzhi Shi, Lihua Lu, Jingwen Leng, Yu Feng, Minyi Guo

首次发表
浏览论文内容

中文总结 AI 辅助

ByteSplat通过融合光栅化内核、硬件感知剪枝和梯度稀疏性通信减少,实现高效分布式3DGS训练,显著降低数据移动并提升训练速度。

中文摘要 AI 辅助

3D高斯溅射(3DGS)能够实现逼真的场景重建,但训练大规模场景需要大量的内存和计算资源。将训练分布到多个GPU上可增加可用内存容量,但其性能严格受数据移动限制。我们识别出两个主要瓶颈:前向和后向光栅化过程中重复的片外内存访问,以及部分高斯梯度的GPU间通信。为解决这些瓶颈,我们提出ByteSplat,一个分布式3DGS训练框架,联合减少GPU内和GPU间的数据移动。首先,ByteSplat将前向光栅化和后向光栅化融合为单个GPU内核,将中间结果保留在片上,以消除冗余的片外传输。其次,为缓解融合光栅化带来的片上存储增加,ByteSplat引入考虑渲染质量和每块共享内存约束的硬件感知剪枝,提升融合执行的整体性能。最后,ByteSplat利用梯度稀疏性消除GPU间通信中的零部分梯度。GPU高效编码器和解码器内核压缩剩余记录,并直接在所有者GPU上聚合接收到的梯度,减少通信量和本地聚合开销。我们在六个数据集上评估ByteSplat。与基线相比,ByteSplat将GPU内片外流量和后向GPU间通信量分别减少63.4%和65.8%。在八块GPU上,ByteSplat在保持重建质量的同时实现高达6.1倍的训练加速。

英文摘要

3D Gaussian Splatting (3DGS) enables photorealistic scene reconstruction, but training large-scale scenes requires substantial memory and computation. Distributing training across multiple GPUs increases available memory capacity, yet its performance is strictly constrained by data movement. We identify two dominant bottlenecks: repeated off-chip memory accesses during forward and backward rasterization, and inter-GPU communication of partial Gaussian gradients. To address these bottlenecks, we present ByteSplat, a distributed 3DGS training framework that jointly reduces intra- and inter-GPU data movement. First, ByteSplat fuses forward rasterization and backward rasterization into a single GPU kernel, retaining intermediate results on-chip to eliminate redundant off-chip transfers. Second, to alleviate the increased on-chip storage by fused rasterization, ByteSplat introduces hardware-aware pruning that considers both rendering quality and per-tile shared-memory constraints, increasing the overall performance of fused execution. Lastly, ByteSplat exploits gradient sparsity to eliminate zero partial gradients from inter-GPU communication. GPU-efficient encoder and decoder kernels compact the remaining records and directly aggregate the received gradients on their owner GPUs, reducing both communication volume and local aggregation overhead. We evaluate ByteSplat across six datasets. Compared with the baseline, ByteSplat reduces intra-GPU off-chip traffic and backward inter-GPU communication volume by 63.4% and 65.8%, respectively. On eight GPUs, ByteSplat achieves up to 6.1$\times$ training speedup while preserving reconstruction quality.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • IEIT SYSTEMS Co., Ltd.(亿道信息系统有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑