发表机构
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics; Institute of Automation, Chinese Academy of Sciences; Department of Computing, Hong Kong Polytechnic University; School of Computer Science, Peking University; Guangdong Laboratory of Artificial Intelligence and Digital Economy; Huawei Consumer Business Group, Huawei; Department of Information Engineering and Computer Science, University of Trento(南京航空航天大学人工智能学院; 中国科学院自动化研究所; 香港理工大学计算学系; 北京大学计算机科学学院; 广东省人工智能与数字经济实验室; 华为消费者业务集团; 特伦托大学信息工程与计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对动态场景重建问题,GrainGS结合分层锚点框架与高斯变形,通过静态预热、停止梯度操作等实现各高斯独立预测时间偏移及规范-残差外观分解,在合成和真实多视图基准测试中展现出高重建质量、实时渲染及紧凑存储等优势。
AI 中文摘要
基于3D高斯点云渲染的动态场景重建需要在细粒度运动建模、结构稳定性和紧凑表示之间取得平衡。现有方法存在冗余原语增长或抑制局部变化运动等问题。为此提出GrainGS,结合分层锚点框架与高斯变形。先进行静态预热建立不变表示,联合训练时通过停止梯度操作,各高斯预测独立时间偏移,还进行规范-残差外观分解。实验表明该方法实现了高重建质量、实时新视图合成和紧凑存储。
英文摘要
Dynamic scene reconstruction with 3D Gaussian Splatting requires a balance between fine-grained motion modeling, structural stability, and compact representation. Existing per-primitive methods provide flexible local deformation but often suffer from redundant primitive growth, while anchor-based methods improve spatial regularity at the cost of suppressing locally varying motion. To address these issues, we present GrainGS, a dynamic Gaussian framework that combines a hierarchical anchor scaffold with per-Gaussian deformation. A static warm-up stage first establishes a time-invariant canonical representation from observations across all timestamps. During joint training, a stop-gradient operation blocks the deformation-mediated gradient pathway to the canonical positions while preserving their direct refinement through the reconstruction objective. Each Gaussian then predicts independent temporal offsets for position, rotation, and scale, enabling detailed local motion within a structurally constrained scaffold. A canonical-residual appearance decomposition further models frame-dependent photometric changes without forcing them into geometric deformation. Experiments on synthetic monocular and real-world multiview benchmarks show that GrainGS achieves high reconstruction quality, real-time novel view synthesis, and compact storage. Under the synthetic benchmark setting, it reaches an average peak signal-to-noise ratio of 36.98 decibels, renders at 435.6 frames per second, and requires 4.67 megabytes of storage.