arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ATSplat:具有自适应令牌扩展的紧凑型前馈3D高斯点渲染

ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion

In Cho, Jeonghwan Cho, Mijin Yoo, Gim Hee Lee, Seon Joo Kim

arXiv 2607.20417首次发表:更新:

发表机构

Yonsei University; National University of Singapore(延世大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有前馈3DGS方法不足,提出ATSplat框架。通过自适应3D令牌恢复自适应分配能力,先提升深度等形成场景支架,再回归高斯并解耦放置,还用模块预测不确定性分数并扩展令牌。实验表明其实现高质量渲染且减少高斯数量,提升效率。

AI 中文摘要

3D高斯点渲染(3DGS)通过在3D中优化自由放置的原语并在重建不足的区域自适应地使其密集化来实现高质量的新视图合成。然而,现有的前馈3DGS方法在很大程度上失去了这种场景自适应能力分配,这些方法通常在输入像素处回归高斯并沿相机光线提升它们。这种像素对齐的公式使得原语的数量和位置取决于图像分辨率和输入视点,而不是场景复杂性,导致密集且通常冗余的高斯集。我们提出了ATSplat,一个前馈3DGS框架,通过自适应3D令牌恢复3DGS优化的自适应分配能力。ATSplat首先将粗略的补丁级深度和相机线索提升到稀疏的3D锚定令牌中,形成场景的紧凑支架。然后,每个令牌通过可学习的3D偏移量回归到局部高斯,将原语放置与输入图像网格解耦。一个自适应令牌扩展模块预测令牌级不确定性分数,由渲染误差图监督,并通过可学习的扩展层选择性地扩展高不确定性令牌。这种从稀疏到自适应的公式使ATSplat能够在具有挑战性的区域集中原语,同时保持紧凑的表示。在两个代表性数据集RealEstate10K和DL3DV上的实验表明,ATSplat实现了最先进的渲染质量,同时与密集的前馈3DGS方法相比,高斯数量减少了5.7倍以上。从12张分辨率为512×960的输入图像中,ATSplat使用单个商用GPU在不到一秒的时间内完成重建,并以1136 FPS(512×960)的速度渲染高质量的新视图,仅使用311K个高斯。

英文摘要

3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely lost in existing feed-forward 3DGS methods, which commonly regress Gaussians at input pixels and lift them along camera rays. Such pixel-aligned formulations make the number and placement of primitives depend on image resolution and input viewpoints rather than scene complexity, resulting in dense and often redundant Gaussian sets. We present ATSplat, a feed-forward 3DGS framework that restores the adaptive allocation capability of 3DGS optimization through Adaptive 3D Tokens. ATSplat first lifts coarse patch-level depth and camera cues into sparse 3D anchor tokens, forming a compact scaffold of the scene. Each token is then regressed into local Gaussians with learnable 3D offsets, decoupling primitive placement from input image grids. An Adaptive Token Expansion module predicts a token-level uncertainty score, supervised by rendering error maps, and selectively expands high-uncertainty tokens through learnable expansion layers. This sparse-to-adaptive formulation enables ATSplat to concentrate primitives in challenging regions while maintaining a compact representation. Experiments on two representative datasets, RealEstate10K and DL3DV, show that ATSplat achieves state-of-the-art rendering quality while reducing the number of Gaussians by more than $5.7\times$ compared with dense feed-forward 3DGS methods. From 12 input images at $512 \times 960$ resolution, ATSplat completes reconstruction in less than a second using a single commercial GPU, and renders high-quality novel views at 1136 FPS ($512 \times 960$) with only 311K Gaussians.

CommentsProject page is at: https://join16.github.io/page-atsplat

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑