arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19893cs.DC

奥丁:基于点的分布式神经渲染的原语级同步

Odin: Primitive-Level Synchronization for Distributed Point-Based Neural Rendering

Zhenxiang Ma, Zeyu He, Yuanzhen Zhou, Zhenyu Yang, Yuchang Zhang, Miao Tao, Rong Fu, Jidong Zhai, Hengjie Li

首次发表
浏览论文内容

中文总结 AI 辅助

研究基于点的神经渲染中的同步问题,提出奥丁分布式训练系统,用原语级同步取代全局屏障,有质量优先和吞吐量优先路径。在多个场景和GPU上实验,提升了吞吐量,减少关键路径等待时间,如在矩阵城市案例中比格伦德尔提高了1.89倍。

中文摘要 AI 辅助

基于点的神经渲染(PBNR)将3D场景表示为显式的、可训练的原语,支撑着高质量重建以及新兴的具身人工智能和世界模型管道。与层结构神经网络不同,PBNR具有原语索引依赖性。大场景需要分布式训练,优化渲染器减少了每个视图的计算量,全局任务或迭代级屏障使同步而非渲染处于关键路径上。我们提出了奥丁,一个分布式PBNR训练系统,它用原语级同步取代全局屏障。其提前调度器利用稳定局部性和阶段顺序识别低冲突重叠窗口,运行时在后续工作观察可变状态之前验证原语发布。奥丁提供了一条保留同步训练可见性的质量优先路径和一条利用重叠和梯度证据仅允许小的、低影响延迟读取的吞吐量优先路径。在四个现有PBNR管道和13个非城市场景的8个GPU上,奥丁平均将吞吐量提高了1.22倍,并隐藏了82%的关键路径等待时间,同时保持重建质量。在一个扩展到64个GPU的矩阵城市混合并行案例研究中,奥丁在不改变渲染器内核、优化器、训练预算或模型容量的情况下,比格伦德尔提高了高达1.89倍的吞吐量。

英文摘要

Point-based neural rendering (PBNR) represents 3D scenes as explicit, trainable primitives and underpins high-quality reconstruction and emerging embodied AI and world-model pipelines. Unlike layer-structured neural networks, PBNR has primitive-indexed dependencies: each view reads and updates only a sparse, view-dependent subset of mutable scene state. As large scenes require distributed training and optimized renderers reduce per-view computation, global task- or iteration-level barriers increasingly place synchronization, rather than rendering, on the critical path. We present Odin, a distributed PBNR training system that replaces global barriers with primitive-level synchronization. Its ahead-of-time scheduler uses stable locality and phase order to identify low-conflict overlap windows, while the runtime validates primitive publication before later work observes mutable state. Odin provides a quality-first path that preserves synchronized-training visibility and a throughput-first path that uses overlap and gradient evidence to admit only small, low-impact delayed reads; structural changes and high-impact cases remain synchronized. Across four existing PBNR pipelines and 13 non-city scenes on 8 GPUs, Odin improves throughput by 1.22 times on average and hides 82% of critical-path wait while preserving reconstruction quality. In a MatrixCity mixed-parallel case study scaling to 64 GPUs, Odin improves throughput over Grendel by up to 1.89 times without changing renderer kernels, optimizers, training budgets, or model capacity.

补充信息

↑