AI 中文总结
提出CDPM,通过几何对齐语义表示与DINO中心特征金字塔,实现跨模态平面图像配准,显著提升精度并降低计算量。
AI 中文摘要
跨模态图像匹配为平面配准建立跨模态的稳定且准确的几何对应关系。现有的语义表示提供了跨模态一致性,但语义相似性并不必然意味着几何对应。同时,细粒度的CNN特征提供了准确的局部细节,但缺乏全局跨模态语义指导以实现稳定的细化。为解决这些问题,我们提出了CDPM,它首先建立几何一致的语义表示,然后在细粒度定位过程中保持其在对应估计中的主导作用。具体来说,我们使用几何一致的跨模态补丁对逐步适应DINOv3,使特征相似性更好地反映真实的跨模态空间对应关系。然后,我们构建了一个以DINO为中心的Feature Pyramid,其中多尺度DINO表示保持稳定的跨模态对应关系,而轻量级CNN分支为精确的局部细化提供辅助结构细节。在三个跨模态数据集上的大量实验证明了CDPM的优越性能。在VIS-IR上,与密集匹配器RoMa相比,CDPM在AUC@3/5/10/20上分别提高了7.36、13.40、13.75和10.42个百分点,并将mACE从5.83像素降低到2.78像素。它还在所有指标上优于RoMa v2,同时所需的FLOPs减少了45.6%。在线演示和数据集可用,代码将在我们的项目页面上发布,网址为https URL。
英文摘要
Cross-modal image matching establishes stable and accurate geometric correspondences across modalities for planar registration. Existing semantic representations provide cross-modal consistency, but semantic similarity does not necessarily imply geometric correspondence. Meanwhile, fine-grained CNN features provide accurate local details but lack global cross-modal semantic guidance for stable refinement. To address these issues, we propose CDPM, which first establishes geometrically consistent semantic representations and then preserves their dominant role in correspondence estimation during fine-grained localization. Specifically, we progressively adapt DINOv3 using geometrically consistent cross-modal patch pairs, enabling feature similarity to better reflect true cross-modal spatial correspondences. We then construct a DINO-Centric Feature Pyramid, where multi-scale DINO representations maintain stable cross-modal correspondences, while a lightweight CNN branch provides auxiliary structural details for precise local refinement. Extensive experiments on three cross-modal datasets demonstrate the superior performance of CDPM. On VIS-IR, compared with the dense matcher RoMa, CDPM improves AUC@3/5/10/20 by 7.36, 13.40, 13.75, and 10.42 percentage points, respectively, and reduces mACE from 5.83 to 2.78 pixels. It also outperforms RoMa v2 across all metrics while requiring 45.6% fewer FLOPs. The online demo and dataset are available, and the code will be released on our project page at https://warren-wzw.github.io/CDPM/.