arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DensiTok:让前馈3D高斯泼溅看到比给定更多的视图

DensiTok: Making Feed-Forward 3D Gaussian Splatting See More Views Than It Is Given

Minhyeok Lee, Jungho Lee, Minseok Kang, Heeseung Choi, Ig-Jae Kim, Sangyoun Lee

arXiv 2610.07958首次发表:更新:

发表机构

Yonsei University; NAVER AI Lab; Korea Institute of Science and Technology (KIST)(延世大学; NAVER AI实验室; 韩国科学技术研究院(KIST))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DensiTok通过直接加密内部几何令牌,使冻结的前馈3DGS主干网络在稀疏视图下重建质量接近密集视图,无需额外图像合成。

AI 中文摘要

前馈3D高斯泼溅(3DGS)通过单次前向传播重建场景,用跨多个场景训练的网络取代了逐场景优化。然而,随着输入图像数量的减少,其质量急剧下降。瓶颈位于重建头之前:从少数未定位的视图,它们读取的内部表示对未观察区域没有证据,导致空洞、漂浮物和模糊。常见的补救措施是以像素形式提供该证据,通过图像或视频生成器合成额外视图并重新编码,这既昂贵又无法保证3D一致性。我们转而加密证据本身。我们提出DensiTok,一个用于预训练前馈3DGS模型的即插即用模块,直接加密其内部几何令牌,使冻结的主干网络表现得如同观察到了比给定更多的视图。DensiTok将这些令牌压缩到紧凑的潜在空间中,在单次流匹配步骤中基于相机几何完成未观察视点的潜在表示,并将其解码回原始重建头所需的令牌。相同的模块设计可以集成到不同的预训练预测器中,同时保持每个主干网络及其重建头冻结。在低维潜在空间中的完成不需要图像合成或额外的编码器通道。在三个预训练主干网络和两个基准测试中,DensiTok持续改善了稀疏视图重建,并弥补了与密集视图重建的大部分差距。

英文摘要

Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across many scenes. Its quality, however, degrades sharply as the number of input images drops. The bottleneck is upstream of the reconstruction heads: from a few unposed views, the internal representation they read carries no evidence for unobserved regions, leaving holes, floaters, and blur. The common remedy supplies that evidence as pixels, synthesizing extra views with an image or video generator and re-encoding them, which is costly and not 3D-consistent by construction. We instead densify the evidence itself. We present DensiTok, a plug-in module for pretrained feed-forward 3DGS models that densifies their internal geometry tokens directly, making a frozen backbone behave as though it had observed many more views than it was given. DensiTok compresses those tokens into a compact latent space, completes the latents of the unobserved viewpoints in a single flow-matching step conditioned on camera geometry, and decodes them back into tokens that the original reconstruction heads. The same module design can be integrated into different pretrained predictors while keeping each backbone and its reconstruction heads frozen. Completion in a low-dimensional latent space requires no image synthesis or additional encoder passes. Across three pretrained backbones and two benchmarks, DensiTok consistently improves sparse-view reconstruction and recovers much of the gap to dense-view reconstruction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑