arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09590cs.CV

TeaMatch:用于2D-3D匹配的可教学跨模态表示学习

TeaMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching

发表机构山东科技大学
查看机构详情
  • Shandong University of Science and Technology(山东科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Chongjian Wang, Junjie Gao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出TeaMatch框架,将可教学性作为跨模态表示学习准则,提升2D-3D匹配鲁棒性,可无缝集成到现有匹配流水线,在相关基准上实现SOTA性能。

中文摘要 AI 辅助

学习图像与点云之间的可靠对应关系是2D-3D匹配的基础。尽管无检测方法已取得进展,但现有方法主要在单一模型内优化匹配,在噪声输入、低重叠、结构模糊等挑战性条件下常难以维持可靠对应。本研究提出TeaMatch,一种将可教学性作为跨模态表示学习准则的新型框架。我们将可教学性定义为:表示在退化输入下能被弱学习器有效恢复的能力,反映其结构一致性与鲁棒性。为此,我们构建一组模拟常见失败模式的特定任务弱学生,在训练集上训练它们模仿教师,并在不相交的元集上评估其可恢复性;随后在对应级别和几何感知约束引导下优化教师,提升学生恢复可靠对应的能力。该框架可无缝集成到现有由粗到细匹配流水线中,无额外推理成本。大量实验表明,TeaMatch提升了匹配鲁棒性,在具挑战性的2D-3D匹配基准上达到了SOTA性能。

英文摘要

Learning reliable correspondences between images and point clouds is fundamental for 2D-3D matching. Despite recent progress in detection-free methods, existing approaches primarily optimize matching within a single model and often struggle to maintain reliable correspondences under challenging conditions such as noisy inputs, low overlap, and ambiguous structures. In this work, we propose TeaMatch, a novel framework that introduces teachability as a criterion for cross-modal representation learning. We define teachability as the ability of a representation to be effectively recovered by weak learners under degraded inputs, reflecting its structural consistency and robustness. To this end, we construct a set of task-specific weak students that simulate common failure modes and train them to imitate the teacher on a training split while evaluating their recoverability on a disjoint meta split. The teacher is then optimized to improve the students' ability to recover reliable correspondences, guided by correspondence-level and geometry-aware constraints. Our framework can be seamlessly integrated into existing coarse-to-fine matching pipelines without additional inference cost. Extensive experiments demonstrate that TeaMatch improves matching robustness and achieves state-of-the-art performance on challenging 2D-3D matching benchmarks.

↑