arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mask 2D-3D:用于图像到点云配准的自适应双掩码自编码器网络

Mask 2D-3D: Adaptive Dual-Masked Autoencoder Network for Image-to-Point Cloud Registration

Zhixin Cheng, Jiacheng Deng, Xiaotian Yin, Baoqun Yin, Richang Hong, Tianzhu Zhang

arXiv 2609.18088首次发表:更新:

发表机构

Hefei University of Technology; University of Science and Technology of China; Institute of Advanced Technology, University of Science and Technology of China(合肥工业大学; 中国科学技术大学; 中国科学技术大学先进技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对图像到点云配准中域差异和错误对应问题,提出基于相似性和强化学习的自适应掩码双MAE框架(ID-MAE),在RGB-D Scenes v2和7-Scenes上取得最先进性能。

AI 中文摘要

无检测的图像到点云配准方法容易因域和模态差异、特征提取器灵敏度有限以及非重叠区域的存在而产生错误对应。掩码自编码器(MAE)在图像和点云的视觉表示方面表现出强大的性能。将这种方法应用于图像到点云配准(一项需要统一特征提取和准确跨模态对应关系的任务)可能是有帮助的。标准MAE的随机掩码可能因相机视野有限而忽略关键区域,降低配准效果。为解决这一问题,我们提出了模态间双MAE框架(ID-MAE),并采用基于相似性的强化学习掩码策略(SRLM),该策略利用跨模态相似性和强化学习自适应地掩码信息丰富的位置,从而缩小模态差距。我们的方法通过在特征提取过程中强制表示一致性来增强跨模态表示学习,从而实现更可靠的2D-3D对应估计。在RGB-D Scenes v2和7-Scenes基准上的实验表明,我们的方法在图像到点云配准方面达到了最先进的性能。

英文摘要

Detection-free methods for image-to-point cloud registration are prone to erroneous correspondences caused by domain and modality discrepancies, limited sensitivity of feature extractors, and the presence of non-overlapping regions. The Masked Autoencoder (MAE) has shown strong performance in visual representation for images and point clouds. It may be helpful to apply this approach to image-to-point cloud registration, a task that requires unified feature extraction and accurate cross-modal correspondences. Standard MAE's random masking may overlook key regions due to limited camera views, reducing registration effectiveness. To address this, we propose the Intermodal Dual-MAE Framework (ID-MAE) with a Similarity-based RL Masking Strategy (SRLM), which adaptively masks informative positions by leveraging cross-modal similarity and reinforcement learning, thus narrowing the modality gap. Our method enhances cross-modal representation learning by enforcing representation consistency during feature extraction, thereby enabling more reliable 2D-3D correspondence estimation. Experiments on RGB-D Scenes v2 and 7-Scenes benchmarks show that our method achieves state-of-the-art performance in image-to-point cloud registration.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑