arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GenRec:明确何处重建、何处生成

GenRec: Knowing Where to Reconstruct and Where to Generate

Ata Çelen, Jaewoo Jung, Federico Tombari, Marc Pollefeys, Sunghwan Hong, Michael Niemeyer, Daniel Barath

arXiv 2608.17832首次发表:更新:

发表机构

KAIST; Google; Microsoft; ETH Zürich(韩国科学技术院; 谷歌公司; 微软公司; 苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GenRec 是一种多视图流匹配模型,通过架构等内置重建与生成划分,在 RealEstate10K 等数据集的单视图外推、两视图插值任务中,兼顾观测区域高重建保真度与未观测区域优感知质量。

AI 中文摘要

从稀疏输入图像生成新视角视图的任务中,生成过程极少是完全重建或完全生成的:在某些源视图中可见的像素具有唯一的正确值,仅受视图相关着色的调制;而在非遮挡区域或超出捕获体积的像素则存在多种合理补全的分布。现有的生成式新视角合成方法将这些不同 regime(区域)统一在单一的均匀损失下,即使通过变形点云或投影深度注入场景几何,也模糊了几何保真度与创造性幻觉之间的界限。我们提出 GenRec,一种多视图流匹配模型,其架构、监督和梯度流中直接内置了重建-生成的划分。在源自源相机和单目深度估计器的观测掩码引导下,流匹配主干网络联合对所有目标视图的 RGB 和场景坐标图进行去噪,同时像素空间细化阶段在观测像素上恢复高频细节;同一掩码控制监督,使回归信号不会污染生成先验。在 RealEstate10K、DL3DV-10K 和 Mip-NeRF 360 数据集上,无论是单视图外推还是两视图插值,GenRec 在观测区域均达到最佳重建保真度,且在未观测区域的感知质量上也超越了纯生成基线,证明了我们方法的有效性。

英文摘要

Generative novel view synthesis from sparse input images is rarely all reconstruction or all generation: pixels visible in some source view have a unique correct value modulated only by view-dependent shading, while pixels in disocclusions or beyond the captured volume admit a distribution of plausible completions. Existing generative novel-view-synthesis methods conflate these regimes under a single uniform loss, blurring the line between geometric fidelity and creative hallucinations even when scene geometry is injected through warped point clouds or projected depth. We introduce GenRec, a multi-view flow matching model that builds the reconstruction--generation split directly into its architecture, supervision, and gradient flow. Guided by an observation mask derived from the source cameras and a monocular depth estimator, a flow matching backbone jointly denoises RGB and scene-coordinate maps across all target views, while a pixel-space refinement stage restores high-frequency detail on observed pixels; the same mask gates supervision so regression signals do not contaminate the generative prior. Across RealEstate10K, DL3DV-10K, and Mip-NeRF~360, in both single-view extrapolation and two-view interpolation, GenRec attains the best reconstruction fidelity in observed regions while also surpassing purely generative baselines on perceptual quality in unobserved ones, showing the effectiveness of our approach.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑