arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RIPE++:仅从正样本对中学习的强化关键点学习

RIPE++: Reinforced Keypoint Learning from Positive Pairs Only

Johannes Künzel, Peter Eisert, Anna Hilsmann

arXiv 2608.19693首次发表:更新:

发表机构

Fraunhofer Heinrich-Hertz-Institute, HHI; Humboldt University Berlin(弗劳恩霍夫海因里希·赫兹研究所; 柏林洪堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出RIPE++,一种仅从正样本对学习的强化关键点学习方法,适配LightGlue后提升了MegaDepth1500的AUC@5,可用于低纹理医学视频等监督有限的场景。

AI 中文摘要

稀疏关键点提取与匹配是几何计算机视觉核心任务的基础,包括运动恢复结构(SfM)、视觉同时定位与地图构建(visual SLAM)、增强现实以及医学图像配准。然而,学习鲁棒的局部特征表示通常需要准确的相机位姿或深度监督,而这些在实际场景中往往不可用。强化学习(RL)近来成为一种有前景的替代方案,仅需判断两张图像是否呈现同一场景的信息。但现有RL公式如RIPE依赖粗略的二元奖励和精心构建的负训练对,限制了训练稳定性和描述子的判别能力。本文重新审视基于RL的关键点学习,提出一种充分利用几何一致性信号的奖励,从单个正样本对中推导奖励和惩罚,无需与负样本对比。这种更丰富的信号提供了足够的监督对比,仅从正图像对中学习判别性检测器和描述子,实现了在极有限监督下的表示学习。此外,我们表明,通过适配LightGlue,相同的RL目标可扩展到匹配阶段,使MegaDepth1500上的AUC@5从56.58提升至59.65,并能从具有部分视觉重叠的图像对中对整个稀疏匹配流水线进行弱监督训练。我们在既定基准上验证了方法,与全监督方法相比展现出有竞争力的结果。进一步表明,该方法甚至可在低纹理医学视频序列上训练,而此类序列通常无法获取相机位姿,且标准SfM流水线常失效。代码和数据可在该https URL获取。

英文摘要

Sparse keypoint extraction and matching underpin core tasks in geometric computer vision, including structure-from-motion, visual SLAM, augmented reality, and medical image registration. Learning robust local feature representations, however, typically requires accurate camera poses or depth supervision, which are often unavailable in real-world settings. Reinforcement learning (RL) has recently emerged as a promising alternative, requiring only the information if two images show the same scene or not. However, existing RL formulations such as RIPE rely on coarse binary rewards and carefully constructed negative training pairs, limiting training stability and descriptor discriminability. In this paper, we revisit RL-based keypoint learning and propose a reward that fully exploits the geometric consistency signal, deriving both reward and penalty from a single positive pair without contrasting against negatives. This richer signal provides sufficient supervisory contrast to learn discriminative detectors and descriptors from positive image pairs alone, enabling representation learning under extremely limited supervision. Furthermore, we show that the same RL objective can be extended to the matching stage by adapting LightGlue, raising AUC@5 on MegaDepth1500 from 56.58 to 59.65 and enabling weakly-supervised training of the full sparse matching pipeline from image pairs with partial visual overlap. We validate our approach on established benchmarks, demonstrating competitive results compared to fully-supervised methods. We further show that the method can be even trained on low texture medical video sequences, where camera poses are usually unavailable and standard SfM pipelines often fail. Code and data are available at https://github.com/fraunhoferhhi/RIPEpp .

CommentsLIMIT@ECCV 2026 (Best Paper Award)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑