发表机构
Neusoft(东软集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Budgeted-GS方法,通过因子化LOD和预算中心训练,将任意3DGS模型转化为多分辨率层级,实现城市规模场景在消费级GPU上的实时高质量渲染。
AI 中文摘要
三维高斯泼溅(3D Gaussian Splatting)在实时渲染中实现了卓越的视觉质量,但在整个城市规模下却难以适用:训练好的模型携带数百万个图元,占用数GB内存,在消费级GPU上实现高质量实时渲染仍然遥不可及。我们提出Budgeted-GS,这是一种事后方法,可将任何训练好的3DGS模型转化为因子化树,即一种多分辨率矩匹配聚合体层级结构。经过几秒钟的构建过程后,一个单一的质量参数即可为每个视图选择适配目标设备内存的细节级别,因此同一城市规模模型可服务于内存容量差异巨大的各类GPU。当需要训练新场景时,同样的理论同样适用:无需先构建完整尺寸模型再压缩,预算中心训练首先测量场景所需图元数量,然后直接以该尺寸训练模型,从而避免优化后续被丢弃图元的浪费。两种方法均基于可测量的容量下限,即从相空间最优传输推导出的预算-误差定律;由近期覆盖定理认证的选择规则决定哪些图元是冗余的。该下限回答了场景实际需要多少图元以及可以安全舍弃多少图元。我们在预注册协议下对13个公共场景验证了该下限,并针对从物体场景到官方城市采集的两种方法进行了实践,在单个消费级GPU上以原生1920x1080分辨率、完整球谐(SH)实时渲染该城市。
英文摘要
3D Gaussian Splatting achieves excellent visual quality with real-time rendering, but at the scale of entire cities it does not fit: a trained model carries millions of primitives and gigabytes of memory, and real-time rendering at high quality on a consumer GPU remains out of reach. We introduce Budgeted-GS, a post-hoc method that turns any trained 3DGS model into a factoring tree, a multi-resolution hierarchy of moment-matched aggregates. After a construction pass of a few seconds, a single quality parameter selects, for each view, the level of detail that fits the memory of the target device, so the same city-scale model serves GPUs with widely different memory capacities. When a new scene is to be trained, the same theory applies: instead of growing a full-sized model and compressing it afterwards, budget-centered training first measures how many primitives the scene needs and then trains the model directly at that size, avoiding the wasted effort of optimizing primitives that are later discarded. Both methods are grounded in a measurable capacity floor, a budget-error law derived from optimal transport in phase space; selection rules certified by recent covering theorems decide which primitives are redundant. The floor answers how many primitives a scene actually needs and how many can safely be given up. We validate the floor on 13 public scenes under a preregistered protocol, and exercise both methods from object scenes to an official city capture, rendering it at native 1920x1080, full SH, in real time on one consumer GPU.
Comments28 pages, 21 figures. Preprint of the EG 2027 submission (paper1075)