发表机构
Shaanxi Key Laboratory for Network Computing and Security Technology, Xi’an University of Technology; College of Computer Science, Sichuan University(陕西理工大学网络计算与安全技术重点实验室; 四川大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究重新思考单目深度嵌入用于广义立体匹配,通过减少分支耦合、构建软约束等方式,将单目信息集成到特征提取和GRU迭代中,还提出相关方法解决边界模糊问题,在多基准测试中达领先性能,提升了泛化能力与准确性。
AI 中文摘要
一般来说,单目方法能捕捉丰富的上下文先验但缺乏几何精度,而立体方法几何精确却在无纹理和遮挡区域表现不佳。一些方法试图结合二者优势,通过对齐单目深度和立体信息来增强立体匹配的泛化能力。然而,建立稳定且可泛化的对齐具有挑战性,不可靠的单目线索会大幅降低性能。本文重新思考单目深度嵌入,减少分支耦合而非扩展网络宽度以防止捷径学习,构建软约束而非硬约束以提高对单目深度误差的容忍度,并将单目信息集成到特征提取和GRU迭代中。具体做法包括融合单目深度图与RGB图像以锐化深度边界感知并抑制匹配模糊性,利用融合图像进行特征提取使上下文特征编码全局几何信息,采用单目深度梯度特征引导视差更新以避免局部振荡,还提出边缘置信度估计方法和边缘感知损失函数来解决数据增强导致的监督视差边界模糊问题。我们的方法在多个标准基准测试中取得了领先性能,展示了出色的泛化能力并提高了准确性。代码可在指定网址获取。
英文摘要
Generally, monocular methods capture rich contextual priors but lack geometric precision, whereas stereo methods are geometrically accurate yet struggle in textureless and occluded regions. Several approaches attempt to combine their strengths to enhance the generalization of stereo matching (SM) by aligning monocular depth with stereo information. However, establishing a stable and generalizable alignment is challenging, and unreliable monocular cues can substantially degrade performance. This paper rethinks monocular depth embedding. First, to prevent shortcut learning, we reduce branch coupling instead of expanding network width. Second, we construct soft constraints instead of hard ones from monocular depth to improve tolerance to monocular depth errors. Based on the principles, we integrate monocular information into both feature extraction and GRU iterations. Specifically, the monocular depth map is fused with the RGB image to sharpen depth boundary perception and suppress matching ambiguities. The fused image is then used for feature extraction, allowing the contextual features to encode global geometric information. Furthermore, the monocular depth gradient feature is employed to guide disparity updates, helping to escape local oscillations. Finally, to address the boundary blurring of supervised disparity caused by data augmentation, we propose an edge confidence estimation method and an edge-aware loss function. Our method achieves state-of-the-art (SOTA) performance on multiple standard benchmarks, demonstrating excellent generalization while improving accuracy. The code is available at https://github.com/linliboabc-maker/stereo-matching-digital.
Comments15 pages, submitted to Pattern Recognition