arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhasorNet:从频率学习结构以实现实时立体匹配

PhasorNet: Learning Structure from Frequency for Real-Time Stereo Matching

Md Raqib Khan, Santosh Kumar Vipparthi, Subrahmanyam Murala

arXiv 2608.29819首次发表:更新:

发表机构

Trinity College Dublin, The University of Dublin; Indian Institute of Technology Ropar(都柏林大学圣三一学院; 印度理工学院罗帕尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PhasorNet是含PAT、GCFRM及EHR损失的轻量立体匹配框架,参数仅530万,在ETH3D基准达SOTA,KITTI泛化性佳,可实现高效实时立体匹配。

AI 中文摘要

精确的立体匹配在病态区域(如精细结构、反光或透明物体)中仍具挑战性,这些区域的外观线索往往模糊或不可靠。为解决该问题,我们提出PhasorNet,这是一个轻量但强大的框架,可通过频域线索提升几何判别能力。其核心是相位增强Transformer(PAT),它将傅里叶导出的相位信息注入注意力机制,生成光度鲁棒、保留结构的特征,在困难区域优先保证结构一致性。此外,我们开发了几何上下文融合细化模块(GCFRM),它结合全分辨率卷积流与轻量注意力流(利用WQA和CDGA模块),可高效保留精细细节和物体边界,且不会产生过多开销。训练进一步通过多尺度边缘引导高误差区域(EHR)损失增强,该损失自适应地将优化聚焦于高误差区域和边缘区域,引导分层代价体细化。仅用530万参数的PhasorNet在极具挑战性的ETH3D基准上达到了最先进性能,同时在KITTI上表现出出色的跨域泛化能力,为精确的实时立体匹配提供了高效且实用的解决方案。

英文摘要

Accurate stereo matching remains challenging in ill-posed regions such as fine structures, reflective, or transparent objects, where appearance cues are often ambiguous or unreliable. To tackle this, we propose PhasorNet, a lightweight yet powerful framework that boosts geometric discrimination via frequency-domain cues. At its core, the Phase-Augmented Transformer (PAT) injects Fourier-derived phase information into the attention mechanism, yielding photometrically robust, structure-preserving features that prioritize structural consistency in difficult areas. Additionally, we develop a Geometry-Context Fusion Refinement Module (GCFRM) that combines a full-resolution convolutional stream with a lightweight attention-based stream (leveraging WQA and CDGA blocks) to efficiently preserve fine details and object boundaries without excessive overhead. Training is further enhanced by a multi-scale Edge-guided High-Error Region (EHR) loss that adaptively focuses optimization on high-error and edge regions, guiding hierarchical cost volume refinement. With only 5.3M parameters, PhasorNet achieves state-of-the-art performance on the challenging ETH3D benchmark while exhibiting excellent cross-domain generalization on KITTI, delivering an efficient and practical solution for accurate real-time stereo matching.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑