发表机构
EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种序列到序列的多视图立体方法,利用全局Transformer和相机感知嵌入联合预测所有视图几何,在多个基准上达到最先进性能。
AI 中文摘要
从多视图图像计算精确的几何结构是计算机视觉中的一个基本问题。最近的馈送(FF)模型联合估计3D几何和相机参数,但即使提供了真实相机参数,它们通常也会因重建模糊性而遭受几何失真。本文研究了已知相机参数下的多视图立体(MVS)问题,并提出了一种连接传统MVS和FF方法的新方法。我们没有将MVS视为仅预测单个参考视图深度的序列到一映射,而是将其重新表述为序列到序列任务,类似于FF模型,联合预测所有输入视图的几何结构。我们引入了一种基于全局Transformer的架构,包含两个显式利用相机诱导先验的组件:射线图嵌入将相机参数注入图像块令牌,使Transformer具有相机感知能力;以及统一的全局成本体积取代传统的每视图成本体积,以联合捕获所有视图的3D结构。在多个公共基准上的大量实验表明,我们的方法达到了最先进的性能,超越了MVS和FF重建基线。
英文摘要
Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera parameters are supplied. In this paper, we study the multi-view stereo (MVS) problem with known camera parameters and propose a novel approach that bridges conventional MVS and FF methods. Rather than casting MVS as a sequence-to-one mapping that predicts depth only for a single reference view, we reformulate it as a sequence-to-sequence task, akin to FF models, that jointly predicts geometry for all input views. We introduce a global transformer-based architecture with two components that explicitly exploit camera-induced priors: ray-map embeddings that inject camera parameters into image patch tokens, making the transformer camera-aware, and a unified global cost volume that replaces conventional per-view cost volumes to jointly capture 3D structure across all views. Extensive experiments on multiple public benchmarks show our approach achieves state-of-the-art performance, surpassing both MVS and FF reconstruction baselines.