发表机构
Central China Normal University; Tsinghua University; Fuzhou University; University of the Chinese Academy of Sciences; Space Engineering University; Beijing Normal University; ByteDance Inc.(华中师范大学; 清华大学; 福州大学; 中国科学院大学; 航天工程大学; 北京师范大学; 字节跳动公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对稀疏视角下3DGS重建的几何模糊与细节缺失问题,提出深度、DINO与扩散协同引导的框架,通过深度补全、视角一致学习和生成先验细化,在多个数据集上显著提升重建质量。
AI 中文摘要
从稀疏输入进行新视角合成对于三维高斯泼溅(3DGS)而言仍具挑战性,原因在于几何模糊、跨视角不一致以及约束不足区域中的细节缺失,导致重建质量下降和渲染不稳定。为解决这些问题,我们提出了 D$^{3}$GS,一种深度- DINO - 扩散引导的稀疏视角高斯重建框架,该框架联合增强了几何与外观。D$^{3}$GS 首先通过基于扩散的补全和 DPT(密集预测 Transformer)细化恢复高分辨率、度量深度图,为高斯初始化提供稳健的几何约束。然后,引入 DINO 引导的视角一致学习,以结构特征增强高斯属性,提升多视角一致性。最后,基于扩散的高斯细化模块将生成先验注入迭代优化策略,增强高斯表示中的高频几何和外观细节。在 DTU、LLFF 和 Mip-NeRF 360 上的实验表明,D$^{3}$GS 相较于强基线取得了持续且显著的改进,消融研究验证了每个组件的有效性和互补作用。
英文摘要
Novel view synthesis from sparse inputs remains challenging for 3D Gaussian Splatting (3DGS) due to ambiguous geometry, cross-view inconsistency, and missing details in under-constrained regions, resulting in degraded reconstruction and unstable rendering. To tackle these issues, we propose D$^{3}$GS, a Depth-DINO-Diffusion guided sparse-view Gaussian reconstruction framework that jointly enhances geometry and appearance. D$^{3}$GS first recovers a high-resolution, metric depth map via diffusion-based completion and DPT (Dense Prediction Transformer) refinement, providing robust Gaussian initialization and geometric constraints. Then, a DINO-guided view-consistent learning is introduced to augment Gaussian attributes with structural features, improving multi-view consistency. Finally, a diffusion-based Gaussian refinement module injects generative priors into an iterative optimization strategy, enhancing high-frequency geometric and appearance details within the Gaussian representation. Experiments on DTU, LLFF, and Mip-NeRF 360 show that D$^{3}$GS achieves consistent and substantial improvements over strong baselines, with ablation studies validating the effectiveness and complementary roles of each component.