arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新审视用于宽基线全向立体视觉的匹配响应和扫描特征体

Revisiting Matching Response and Swept Feature Volumes for Wide-baseline Omnidirectional Stereo

Seungjin Jeon, Jongwoo Lim, Changhee Won

arXiv 2607.11097首次发表:更新:

发表机构

UVify Corporate Affiliated Research Institute; Seoul National University(UVify公司附属研究所; 首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对宽基线全向立体视觉中模糊匹配问题,提出训练策略,通过重新解释匹配响应获取置信度信号,直接惩罚模糊响应。还引入扫描特征体重采样,联合学习提升置信度估计与表面法线预测,增强深度一致性且保持部署实用性。

AI 中文摘要

在本文中,我们针对宽基线设置中频繁出现的模糊匹配,提出了一种用于全向立体视觉中置信度估计的训练策略。通过重新解释3D编码器-解码器模块产生的匹配响应,我们表明其期望值提供了内在的置信度信号。在此基础上,我们的方法直接对模糊响应进行惩罚,无需辅助头、多遍推理或额外模块,从而实现更高效和通用的预测。除了置信度,我们还引入了扫描特征体重采样,其中3D CNN产生的响应特征使用回归的正匹配索引进行重采样,然后由2D CNN处理以预测诸如表面法线等元信息。这种联合学习引入了辅助几何正则化,并通过在响应聚合阶段利用额外的上下文线索提高了深度一致性。实验结果表明,我们的方法在保持自主移动应用的部署实用性的同时,增强了置信度估计和表面法线预测。

英文摘要

In this paper, we propose a training strategy for confidence estimation in omnidirectional stereo, targeting the ambiguous matches that frequently occur in wide-baseline setups. Reinterpreting the matching responses produced by the 3D encoder decoder block, we show that their expectation values provide intrinsic confidence signals. Building on this, our method directly penalizes ambiguous responses without auxiliary heads, multi-pass inference, or additional modules, resulting in more efficient and generalized predictions. Beyond confidence, we introduce swept feature volume resampling, where response features produced by 3D CNNs are resampled using regressed positive matching indices and then processed by 2D CNNs to predict meta-information such as surface normals. This joint learning introduces auxiliary geometric regularization and improves depth coherence by leveraging additional contextual cues during response aggregation stage. Experimental results demonstrate that our approach enhances both confidence estimation and surface normal prediction while maintaining deployment practicality for autonomous mobility applications.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑