发表机构
Nanyang Technological University; Harvard University(南洋理工大学; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AVSplat通过辅助视图预条件和占用率引导的体素融合,解决稠密视角下前馈3D高斯泼溅的性能退化问题,实现正向视角缩放。
AI 中文摘要
无位姿前馈3D高斯泼溅(3D Gaussian Splatting)能够从未标定的多视角图像中实现新视角合成。尽管增加视角理应提升性能,但现有方法在稠密视角输入下常常出现性能退化,这是因为全局聚合将注意力分散到大量令牌上,且朴素的体素融合将众多高斯平均成过度平滑的表示。我们提出AVSplat,一个将额外视角转化为聚合和表示两方面可靠信号的框架。在全局注意力之前,每个视角与一组按相关性和多样性选取的小规模辅助视图(Assist Views)进行单次轻量交互,缓存的特征提供了聚焦的场景上下文,从而稳定对应关系。在表示方面,我们采用自适应的温度感知体素融合,在高占用率下锐化归属,并由占用率和点置信度引导。关键在于,AVSplat恢复了正向视角缩放,即随着输入视角增加,性能保持稳定或提升,而非在稠密视角区域退化。消融实验表明,辅助视图预条件(Assist View Preconditioning)主要负责防止稠密视角退化,而占用率引导的体素融合(Occupancy-guided Voxel Fusion)贡献了大部分单点图像质量增益。
英文摘要
Pose-free feed-forward 3D Gaussian Splatting enables novel view synthesis from uncalibrated multi-view images. Although more views should improve performance, existing methods often degrade with dense-view inputs because global aggregation spreads attention over many tokens, and naive voxel fusion averages many Gaussians into overly smooth representations. We present AVSplat, a framework that turns additional views into reliable signals for both aggregation and representation. Before global attention, each view performs a single lightweight interaction with a small set of Assist Views chosen for relevance and diversity, and the cached features provide a focused scene context that stabilizes correspondence. For representation, we use adaptive temperature-aware voxel fusion that sharpens attribution under high occupancy, guided by occupancy and point confidence. Crucially, AVSplat restores positive view scaling where performance remains stable or improves as more input views are added, instead of degrading in the dense-view regime. Ablations show that Assist View Preconditioning is primarily responsible for preventing dense-view degradation, while Occupancy-guided Voxel Fusion contributes most of the single-point image-quality gains.