发表机构
Dongguk University; dot(东国大学; 42dot)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
M2Depth框架通过双向互细化策略耦合深度基础模型与级联多视图立体匹配流程,引入先验引导的代价体细化机制,在标准基准上优于现有多视图立体匹配方法,且对稀疏视图设置泛化能力出色。
AI 中文摘要
基于深度学习的多视图立体匹配(Multi-View Stereo, MVS)已取得显著进展,但往往对未见过的场景泛化能力较差,尤其是在遮挡区域或视图重叠有限的区域。为缓解这一问题,近期的方法将深度基础模型(Depth Foundation Models, DFMs)整合到MVS流程中以提供单目深度先验。然而,现有方法通常依赖静态的单向融合方案,无法充分利用两种模态的互补优势。我们提出一种新颖框架,通过双向互细化策略将DFM与级联MVS流程紧密耦合,克服这一局限。我们的方法利用MVS深度解决单目预测中的尺度歧义,而单目深度则反过来增强MVS估计的结构完整性和细粒度细节。此外,我们引入先验引导的代价体细化机制,通过基于注意力的融合和离散深度bin有效整合多视图与单目信息,从而促进局部几何一致性。大量实验表明,我们的方法在标准基准上优于当前最优的MVS方法,生成更完整、泛化能力更强且边界清晰的深度图。此外,尽管并非专为稀疏视图设置设计,我们的框架仍表现出出色的泛化能力,甚至可与专用的稀疏视图方法相媲美,同时保持更优的精度-效率权衡。
英文摘要
Deep learning-based Multi-View Stereo (MVS) has advanced significantly but often generalizes poorly to unseen scenes, particularly in occluded areas or regions with limited view overlap. To mitigate this, recent approaches integrate Depth Foundation Models (DFMs) into MVS pipelines to provide monocular depth priors. However, existing methods typically rely on a static, one-way fusion scheme, which fails to fully exploit the complementary strengths of both modalities. We propose a novel framework that overcomes this limitation by tightly coupling a DFM with a cascade MVS pipeline through a bidirectional mutual refinement strategy. Our method leverages MVS depth to resolve the scale ambiguity in monocular predictions, while the monocular depth, in turn, enhances the structural completeness and fine-grained detail of the MVS estimate. Furthermore, we introduce a prior-guided cost volume refinement mechanism that effectively integrates multi-view and monocular information via attention-based fusion and discretized depth bins, thereby promoting local geometric consistency. Extensive experiments demonstrate that our method outperforms state-of-the-art MVS approaches on standard benchmarks, producing more complete and generalizable depth maps with sharp boundaries. Furthermore, although not explicitly designed for sparse-view settings, our framework generalizes remarkably well, competing favorably with even dedicated sparse-view methods while maintaining a superior accuracy-efficiency trade-off.