发表机构
Shanghai Jiaotong University; Ningbo Institute of Digital Twin, Eastern Institute of Technology(上海交通大学; 宁波东方理工大学数字孪生研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有立体匹配方法难以保留细粒度几何细节的问题,提出基于扩散的StereoDiffuer框架,通过显著性注意力感知模块提取几何线索,结合迭代去噪扩散过程优化视差,在Scene Flow与KITTI基准上表现具竞争力。
AI 中文摘要
随着深度神经网络的发展,立体匹配得到的视差图质量稳步提升,但现有立体匹配方法仍难以保留细粒度几何细节,导致挑战性区域出现边缘模糊与预测过平滑问题。为解决这些局限,我们提出StereoDiffuer,一种基于迭代扩散的立体匹配框架,该框架显式建模几何细节并渐进式优化视差估计。框架集成显著性注意力感知(Saliency Attention Perception,SAP)模块,用于提取显著几何线索,包括物体边界、细结构与锐边;将置信度引导的SAP特征与初始视差估计结合,用于条件化迭代去噪扩散过程,以修正残余视差误差并恢复代价体正则化与上采样过程中被抑制的几何细节。在Scene Flow与KITTI基准上的实验结果表明,所提框架有效且相较于对比立体匹配方法具备竞争力。
英文摘要
With the advance of deep neural networks, the quality of disparity maps obtained through stereo matching has steadily improved. However, existing stereo matching methods still struggle to preserve fine-grained geometric details, resulting in blurred edges and over-smoothed predictions in challenging regions. To address these limitations, we propose StereoDiffuer, an iterative diffusion-based stereo matching framework that explicitly models geometric details and progressively refines disparity estimates. The framework incorporates a Saliency Attention Perception (SAP) module to extract salient geometric cues, including object boundaries, thin structures, and sharp edges. Confidence-guided SAP features are combined with the initial disparity estimate to condition an iterative denoising diffusion process, which corrects residual disparity errors and restores geometric details suppressed during cost-volume regularization and upsampling. Experimental results on the Scene Flow and KITTI benchmarks demonstrate the effectiveness of the proposed framework and its competitive performance relative to the compared stereo matching methods.
Comments17 pages, 9 figures, and 11 tables. Accepted by Signal Processing: Image Communication