面向功能性灵巧抓取的多关键点可供性表示
Multi-Keypoint Affordance Representation for Functional Dexterous Grasping
- School of Artificial Intelligence and Robotics, Hunan University(人工智能与机器人学院,湖南大学)
- College of Computer Science and Electronic Engineering, Hunan University(计算机科学与电子工程学院,湖南大学)
- National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University(机器人视觉感知与控制技术国家工程研究中心,湖南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对功能性灵巧抓取中视觉感知与操作脱节的问题,提出多关键点可供性表示CMKA与抓取矩阵变换KGT方法,通过定位接触点直接约束抓取姿态,显著提升了抓取精度与泛化能力。
AI中文摘要:
功能性灵巧抓取需要精确的手物交互,超越了简单的抓握。现有的基于可供性的方法主要预测粗略的交互区域,无法直接约束抓取姿态,导致视觉感知与操作之间的脱节。为了解决这个问题,我们提出了一种用于功能性灵巧抓取的多关键点可供性表示,通过定位功能性接触点直接编码任务驱动的抓取配置。我们的方法引入了Contact-guided Multi-Keypoint Affordance (CMKA),利用人类抓取经验图像进行弱监督,并结合Large Vision Models进行精细的可供性特征提取,在避免手动关键点标注的同时实现了泛化。此外,我们提出了一种基于关键点的抓取矩阵变换(Keypoint-based Grasp matrix Transformation, KGT)方法,确保手部关键点与物体接触点之间的空间一致性,从而在视觉感知与灵巧抓取动作之间提供直接联系。在公开的真实世界FAH数据集、IsaacGym仿真以及具有挑战性的机器人任务上的实验表明,我们的方法显著提高了可供性定位精度、抓取一致性以及对未见工具和任务的泛化能力,弥合了视觉可供性学习与灵巧机器人操作之间的差距。源代码和演示视频公开于 https://github.com/PopeyePxx/MKA。
英文摘要:
Functional dexterous grasping requires precise hand-object interaction, going beyond simple gripping. Existing affordance-based methods primarily predict coarse interaction regions and cannot directly constrain the grasping posture, leading to a disconnection between visual perception and manipulation. To address this issue, we propose a multi-keypoint affordance representation for functional dexterous grasping, which directly encodes task-driven grasp configurations by localizing functional contact points. Our method introduces Contact-guided Multi-Keypoint Affordance (CMKA), leveraging human grasping experience images for weak supervision combined with Large Vision Models for fine affordance feature extraction, achieving generalization while avoiding manual keypoint annotations. Additionally, we present a Keypoint-based Grasp matrix Transformation (KGT) method, ensuring spatial consistency between hand keypoints and object contact points, thus providing a direct link between visual perception and dexterous grasping actions. Experiments on public real-world FAH datasets, IsaacGym simulation, and challenging robotic tasks demonstrate that our method significantly improves affordance localization accuracy, grasp consistency, and generalization to unseen tools and tasks, bridging the gap between visual affordance learning and dexterous robotic manipulation. The source code and demo videos are publicly available at https://github.com/PopeyePxx/MKA.