VoxelTTO:体素对齐的前馈三维高斯泼溅与测试时优化
VoxelTTO: Voxel-Aligned Feed-Forward 3D Gaussian Splatting with Test-Time Optimization
浏览论文内容
中文总结 AI 辅助
VoxelTTO提出体素对齐的前馈3DGS框架,通过全局体素表示、测试时优化和实体体积渲染,提升新视角合成与相机姿态估计精度。
中文摘要 AI 辅助
近期前馈三维高斯泼溅(3DGS)方法通常回归像素对齐的高斯原语,常导致过度重叠和伪影,而预测相机位姿的不准确可能导致新视角合成(NVS)中的错位。我们提出VoxelTTO,一种前馈框架,可从任意数量的图像和可选相机参数重建几何精确的3DGS场景。VoxelTTO将密集图像特征聚合为全局体素表示,并从体素特征解码高斯,打破了像素到高斯的对应关系。为利用已知相机参数同时保持预训练视觉基础模型(VFM)参数冻结,我们引入测试时优化(TTO),通过姿态监督适配轻量级LoRA模块。我们进一步在训练和推理期间用随机实体体积渲染替代普通3DGS光栅化,提高几何保真度。训练仅更新体素对齐的高斯重建模块,需80 GPU小时。在Replica、Tanks and Temples和DTU上的实验表明,相较于先前方法,在RGB-D NVS和相机姿态估计方面均有改进。
英文摘要
Recent feed-forward 3D Gaussian Splatting (3DGS) methods typically regress pixel-aligned Gaussian primitives, often causing excessive overlap and artifacts, while inaccuracies in predicted camera poses can lead to misalignment in novel-view synthesis (NVS). We present VoxelTTO, a feed-forward framework for reconstructing geometrically accurate 3DGS scenes from an arbitrary number of images and optional camera parameters. VoxelTTO aggregates dense image features into a global voxel representation and decodes Gaussians from voxel features, breaking the pixel-to-Gaussian correspondence. To exploit known camera parameters while keeping the pretrained visual foundation model (VFM) parameters frozen, we introduce test-time optimization (TTO) that adapts lightweight LoRA modules using pose supervision. We further replace vanilla 3DGS rasterization with stochastic solid volume rendering during training and inference, improving geometric fidelity. Training updates only the voxel-aligned Gaussian reconstruction modules, requiring 80 GPU hours. Experiments on Replica, Tanks and Temples, and DTU demonstrate improved RGB-D NVS and camera-pose estimation relative to prior methods.
发表机构
- East China University of Science and Technology(华东理工大学)
- Shanghai Open University(上海开放大学)
机构由 AI 辅助整理,请以论文原文为准。