发表机构
Korea University; NCSOFT(高丽大学; NCSOFT)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AIMS通过锚点集成和可学习积分器,在固定全局视图预算下融合多视图信息,实现可扩展的新视角渲染,在RealEstate10K和ScanNet上取得高质量与高效率的权衡。
AI 中文摘要
前馈式新视角合成方法能够从带位姿的多视图输入中实现强泛化,但将其扩展到大规模输入视图集仍然具有挑战性。基于Transformer的方法联合处理所有输入视图的令牌,随着视图数量增长,计算和内存需求迅速增加,而简单的视图子采样则会丢弃可能有用的观测。我们提出了锚点集成多视图合成(AIMS),这是一个可扩展框架,将可用观测数量与全局合成模型处理的视图数量解耦。AIMS使用最远点采样选择一组固定且空间分布的锚点视图,将每个锚点附近的观测分组,并使用轻量级可学习积分器将其信息融合为增强的锚点表示。这使得额外的观测能够为合成做出贡献,同时保持下游全局视图预算固定。在RealEstate10K和ScanNet上的评估表明,与基于Transformer和高斯的方法相比,AIMS实现了良好的质量-效率权衡。AIMS在两个数据集上分别达到29.41 dB和17.73 dB的PSNR,每视角渲染平均耗时7.24毫秒。
英文摘要
Feed-forward novel view synthesis methods achieve strong generalization from posed multi-view inputs, but scaling them to large input view sets remains challenging. Transformer-based approaches that jointly process all input-view tokens incur rapidly increasing computation and memory as the number of views grows, while simple view subsampling discards potentially useful observations. We introduce Anchor-Integrated Multi-View Synthesis (AIMS), a scalable framework that decouples the number of available observations from the number of views processed by the global synthesis model. AIMS selects a fixed set of spatially distributed anchor views using farthest point sampling, groups nearby observations around each anchor, and uses a lightweight learnable integrator to fuse their information into enriched anchor representations. This allows additional observations to contribute to synthesis while keeping the downstream global view budget fixed. Evaluations on RealEstate10K and ScanNet demonstrate a favorable quality--efficiency trade-off against transformer-based and Gaussian-based baselines. AIMS achieves 29.41 dB and 17.73 dB PSNR on the two datasets, respectively, with rendering averaging 7.24 ms per view.