arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从斑点(Splat)到硅(Silicon):重新思考3DGS的计算效率

From Splats to Silicon: Rethinking Computational Efficiency of 3DGS

Minnan Pei, Qiwei Dong, Yihan Zhou, Gang Li, Yuchen Zhu, Wenju Zhao, Zhongtian Long, Siting Wang, Peisong Wang, Jian Cheng

arXiv 2609.06157首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Nanjing University; Eindhoven University of Technology; Huazhong University of Science and Technology(中国科学院自动化研究所; 中国科学院大学; 南京大学; 埃因霍温理工大学; 华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出以工作负载为中心的框架,结合复现测量与GPU剖析,分析3DGS渲染与更新路径上的效率瓶颈,揭示系统收益取决于工作负载削减、粒度匹配及数据移动成本,并指出未来系统设计方向。

AI 中文摘要

3D高斯溅射(3DGS)通过显式基元表示场景并支持实时新视角合成,然而其系统效率在不同场景、视角、渲染路径和平台约束下差异显著。现有研究通过表示与算法设计、GPU运行时优化以及架构支持来追求效率,但所报告的性能提升对应于渲染和更新路径上的不同环节。将这些指标与端到端系统收益关联起来,需要追踪每次优化如何改变高斯选择、屏幕空间工作量、数据移动以及阶段或帧时间。因此,我们采用以工作负载为中心的框架,将表示与算法研究、GPU运行时和硬件架构联系起来,并识别重复出现的工作负载模式。我们通过复现测量和对选定实现进行受控GPU剖析来补充文献分析,将工作负载计数与阶段时间和内存流量相关联。综合这些比较表明,系统收益取决于工作负载减少能否传导至下游执行、粒度与每个阶段的匹配程度,以及数据传输、同步、缓存结果、梯度和优化器数据的成本。基于这些发现,我们讨论了在渲染质量约束下更一致的评估方法,并指出了未来系统设计的关键方向。

英文摘要

3D Gaussian splatting (3DGS) represents scenes with explicit primitives and supports real-time novel-view synthesis, yet its system efficiency varies substantially across scenes, viewpoints, rendering paths, and platform constraints. Existing studies pursue efficiency through representation and algorithm design, GPU runtime optimization, and architectural support, but their reported gains correspond to different points along the rendering and update paths. Connecting these indicators to end-to-end system benefit requires tracing how each optimization changes Gaussian selection, screen-space work, data movement, and stage or frame time. We therefore use a workload-centric framework to connect representation and algorithm research, GPU runtimes, and hardware architectures and to identify recurring workload patterns. We complement literature analysis with reproduced measurements and controlled GPU profiling of selected implementations, relating workload counts to stage time and memory traffic. Together, these comparisons show that system gains depend on workload reductions reaching downstream execution, granularity matching each stage, and the cost of data transfers, synchronization, and cached results, gradients, and optimizer data. Building on these findings, we discuss more consistent evaluation under rendering-quality constraints and identify key directions for future system design.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑