发表机构
School of Computer Science and Engineering, Beihang University; University of California, Berkeley; University of Southern California; Google; State Key Laboratory of Intelligent Manufacturing Equipment and Technology, Huazhong University of Science and Technology; Tencent; GigaAI(北京航空航天大学计算机科学与工程学院; 加州大学伯克利分校; 南加州大学; 谷歌公司; 华中科技大学智能制造装备与技术国家重点实验室; 腾讯公司; 极佳科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对立体匹配问题,提出先验引导的生成框架StereoFlow,通过两阶段渐进级联匹配网络、像素扩散变压器StereoDiT和过渡流匹配实现高效优化,在多基准测试中取得多个最新结果,提升了立体匹配效果。
AI 中文摘要
立体匹配是3D重建中的一项基本任务。尽管取得了显著进展,但主流范式将立体匹配表述为确定性回归问题,将多模态分布建模简化为单点估计。这种表述存在回归到均值偏差,在模糊区域经常遇到困难。相比之下,我们引入了一个先验引导的生成框架,将确定性匹配回归和生成分布建模整合在一个互补的公式中。在此基础上,我们通过三个关键组件引入了StereoFlow:一个两阶段渐进级联匹配网络,一个具有频率解耦架构的像素扩散变压器(称为StereoDiT),以及一个用于高效优化的几步流匹配目标(称为过渡流匹配)。总之,StereoFlow在不适定、不连续区域和零样本泛化下实现了强大的几何一致性和丰富的细粒度细节。大量实验表明,所提出的StereoFlow在多个基准测试中建立了多个最新结果,包括场景流、KITTI、ETH3D和米德尔伯里。
英文摘要
Stereo matching is a fundamental task in 3D reconstruction. Despite remarkable advances, the prevailing paradigms formulate stereo matching as a deterministic regression problem, collapsing the multimodal distribution modeling into a single-point estimation. This formulation suffers from a regression-to-mean bias, frequently struggling with ambiguous regions. In contrast, we introduce a prior-guided generative framework that integrates deterministic matching regression and generative distribution modeling within a complementary formulation. Built upon this formulation, we introduce StereoFlow through three key components: (i) a two-stage progressive cascade matching network that progressively produces multi-resolution stereo conditions with complementary matching cues; (ii) a pixel diffusion transformer (termed StereoDiT) with a frequency-decoupled architecture for modeling correspondence ambiguity; (iii) a few-step flow matching objective (termed Transition Flow Matching) for efficient optimization. In summary, \textsc{\textbf{StereoFlow}} achieves strong geometric consistency and rich fine-grained details in ill-posed, discontinuous regions and under zero-shot generalization. Extensive experiments demonstrate that the proposed StereoFlow establishes multiple state-of-the-art results across benchmarks, including Scene Flow, KITTI, ETH3D, and Middlebury.
Comments10 pages, 6 figures, submitted to TVCG