arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35573cs.CV

少即是多:用于高效新视角合成的遗传帧选择

Less Is More: Genetic Frame Selection for Efficient Novel View Synthesis

Diego E. Farchione, Ramzi Idoughi, Alberto Jaspe-Villanueva, Peter Wonka

首次发表
浏览论文内容

中文总结 AI 辅助

针对前馈新视角合成中冗余帧降低效率与质量的问题,提出免渲染的遗传帧选择器,基于覆盖度、冗余度和清晰度评分,在多个数据集上优于基线并降低选择成本。

中文摘要 AI 辅助

前馈式新视角合成通过单次前向传播从多张输入图像重建场景,但更多视角并不一定能提升性能:冗余或选择不当的帧会增加计算成本,并可能降低重建质量。我们解决的问题是,从已捕获的序列中,选择一个固定大小的输入视图子集,该子集对重建指定目标视点最具信息量。我们提出了一种免渲染的视图选择器,它基于三个互补标准对候选帧进行评分:目标视图覆盖度(以代表目标的观测帧为衡量基准)、与已选视图的冗余度以及图像清晰度。一个轻量级评分网络随后选择最具信息量的帧,在推理时无需渲染、重建或逐场景优化。为了训练该选择器,我们蒸馏了一种昂贵的离线搜索过程,其中遗传算法通过直接优化训练场景上的重建性能来识别高质量子集。选择器仅从几何和图像级特征学习复现这些选择。在六个数据集和多种输入预算下,我们的方法始终优于基于几何和基于重建感知的视图选择基线,同时选择成本显著低于基于重建的替代方案。此外,精心选择的子集可以优于从完整输入序列进行的前馈重建。学习到的选择器可泛化到多种重建范式(前馈、3D高斯泼溅和NeRF)、面向目标的物体重建以及跨采集设置(目标视图来自单独采集过程)。更广泛地说,我们的结果表明,显式地考虑目标相关性和视图间冗余是高效场景重建的基本因素。

英文摘要

Feed-forward novel view synthesis reconstructs a scene from many input images in a single forward pass, yet more views do not necessarily improve performance: redundant or poorly chosen frames increase computational cost and may degrade reconstruction quality. We address the problem of selecting, from an already captured sequence, a fixed-size subset of input views that is most informative for reconstructing specified target viewpoints. We propose a render-free view selector that scores candidate frames based on three complementary criteria: target-view coverage, measured against observed frames that stand in for the targets, redundancy with previously selected views, and image sharpness. A lightweight scoring network then selects the most informative frames without rendering, reconstruction, or per-scene optimization at inference time. To train the selector, we distill an expensive offline search procedure in which a genetic algorithm identifies high-quality subsets by directly optimizing reconstruction performance on training scenes. The selector learns to reproduce these choices from geometric and image-level features alone. Across six datasets and multiple input budgets, our method consistently outperforms both geometric and reconstruction-aware view-selection baselines while incurring significantly lower selection costs than reconstruction-based alternatives. Moreover, carefully selected subsets can outperform feed-forward reconstruction from the full input sequence. The learned selector generalizes across diverse reconstruction paradigms (feed-forward, 3D Gaussian Splatting, and NeRF), to object-targeted reconstruction and to a cross-capture setting in which the target views come from a separate acquisition pass. More broadly, our results indicate that explicitly reasoning about target relevance and inter-view redundancy is a fundamental factor in efficient scene reconstruction.

发表机构

  • King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑