AI 中文总结
本研究提出一组捕捉关键点自一致性的几何特征,训练逻辑回归分类器检测位姿估计失败,性能优于仅依赖关键点不确定性的基于置信度的方法。
AI 中文摘要
位姿估计的一种常见方法是预测图像中的物体关键点,随后使用 Perspective-n-Point 算法计算物体相对于相机的旋转和平移。虽然旋转能保持物体形状,但基于关键点的位姿估计方法常忽略该特性,这类方法中关键点通常相互独立预测。由于不精确的关键点预测会负面影响位姿估计精度,也限制了其在下游任务中的可靠性。本研究探索能否仅通过检查2D关键点间的空间位置来识别此类不精确位姿估计。我们提出一组手工设计的几何特征,用于捕捉关键点预测的自一致性,包括成对距离、重投影一致性,以及渲染和掩码一致性。尽管方法简单,基于这些特征训练的逻辑回归分类器可可靠检测位姿估计失败,性能优于仅依赖关键点不确定性的 conformal keypoint predictions 等基于置信度的方法。
英文摘要
One common approach to pose estimation involves predicting object keypoints in an image, followed by using Perspective-n-Point algorithms to compute the object's rotation and translation relative to the camera. While rotations preserve object shapes, this property is often neglected in keypoint-based pose estimation methods, where keypoints are typically predicted independently from each other. As imprecise keypoint predictions negatively affects pose estimation accuracy, it also limits its reliability in downstream tasks. In this work, we explore whether such inaccurate pose estimates can be identified by simply examining spatial locations between 2D keypoints. We propose a set of hand-crafted geometric features that capture the self-consistency of keypoint predictions, including pairwise distances, reprojection consistency, as well as render and mask consistency. Despite its simplicity, a logistic regression classifier trained on these features reliably detects pose estimation failures, outperforming confidence-based approaches like conformal keypoint predictions that rely solely on keypoint uncertainty.