TriCons-Pose:用于类别级目标姿态估计的三角形不变几何一致性学习
TriCons-Pose: Triangle-Invariant Geometric Consistency Learning for Category-Level Object Pose Estimation
- School of Cyberspace, Hangzhou Dianzi University(杭州电子科技大学网络空间学院)
- School of Communication Engineering, Hangzhou Dianzi University(杭州电子科技大学通信工程学院)
- Université Sorbonne Paris Nord, L2TI, UR 3043(巴黎北索邦大学,L2TI实验室,UR 3043研究单位)
- Université Paris-Saclay, CentraleSupélec, CVN(巴黎萨克雷大学,中央理工-高等电力学院,CVN实验室)
- School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究类别级目标姿态估计问题,提出TriCons-Pose方法,设计SCKD识别鲁棒关键点,PIGA增强关键点表示,通过标准目标函数及几何一致性损失优化框架,实验证明该方法有效。
AI中文摘要:
类别级目标姿态估计在学术界和工业界都是一项关键但具有挑战性的任务,通过基于关键点的对应范式取得了显著成功。然而,大多数现有方法越来越依赖更强的特征学习,而忽略了建立的对应关系在不同扰动下是否几何稳定。这通常导致在类内形状变化和遮挡下姿态恢复脆弱。为应对这一挑战,我们开发了一种用于类别级目标姿态估计的新型三角形不变几何一致性学习(TriCons-Pose),以锚定稳定关键点并聚合姿态不变线索,产生可靠的规范映射和准确的姿态估计。具体而言,设计了一个结构一致关键点检测器(SCKD),通过归一化成对距离匹配来强制跨视图结构一致性,以识别鲁棒关键点。此外,我们提出了一个姿态不变几何聚合器(PIGA),通过将基于三角形的姿态不变描述符注入局部到全局注意力机制来增强关键点表示。所提出的框架使用标准目标函数进行优化,同时纳入额外的几何一致性损失。在REAL275、CAMERA25和HouseCat6D数据集上的大量实验证明了该方法的有效性。
英文摘要:
Category-level object pose estimation is a crucial yet challenging task in both academia and industry, and has achieved remarkable success by leveraging keypoint-based correspondence paradigms. However, most existing methods increasingly rely on stronger feature learning while overlooking whether the established correspondences are geometrically stable across diverse perturbations. This often results in fragile pose recovery under intra-class shape variations and occlusions. To tackle this challenge, we develop a novel Triangle-Invariant Geometric Consistency Learning for Category-Level Object Pose Estimation (TriCons-Pose) to anchor stable keypoints and aggregate pose-invariant cues, yielding reliable canonical mapping and accurate pose estimation. Specifically, a Structure-Consistent Keypoint Detector (SCKD) is designed to identify robust keypoints by enforcing cross-view structural consistency via normalized pairwise distance matching. Moreover, we propose a Pose-Invariant Geometric Aggregator (PIGA) to augment keypoint representations by injecting triangle-based pose-invariant descriptors into a local-to-global attention mechanism. The proposed framework is optimized using standard objective functions while incorporating an additional geometry consistency loss. Extensive experiments on REAL275, CAMERA25, and HouseCat6D datasets demonstrate the effectiveness of the proposed approach.