发表机构
Southern University of Science and Technology; KTH Royal Institute of Technology; Shanghai Jiao Tong University; Rysen Robotics(南方科技大学; 瑞典皇家理工学院; 上海交通大学; 睿森机器人)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出 FunCo-Grasp,通过功能部分对齐和规范框架对齐建立跨具身功能对应,使扩散模型学习可迁移抓取知识,在仿真和真实实验中显著提升对未见手的抓取成功率。
AI 中文摘要
跨具身灵巧抓取生成仍然具有挑战性,因为机器手在几何、拓扑和运动学上存在显著差异。现有方法通常缺乏在抓取中扮演相似功能角色的结构不同手部区域之间的显式对应关系,我们将这一概念称为功能对应。因此,它们的模型往往学习手部特定的交互模式,而非可迁移的抓取知识,从而限制了对未见过的泛化。为解决这一局限,我们提出了 FunCo-Grasp,它在异构手部具身之间建立了功能对应关系。具体而言,功能部分对齐通过根据抓取角色将物理链接映射到共享功能部分,将每只手对齐到规范功能模式,而规范框架对齐则将这些部分在规范局部框架中表达。这两种对齐为部分间和手-物体交互提供了一致的表示,使模型能够学习跨手的可迁移抓取知识。以对齐的手部表示和物体几何为条件,扩散模型生成功能部分的目标空间排列,然后将其转换为可执行的关节配置。将 FunCo-Grasp 适应到未见手部仅需其几何和运动学模型以及一次性的轻量功能标注,无需目标手抓取数据、微调或学习的重定向。在过滤后的 CMapDataset 上对保留物体的仿真中,我们在三只已见手上实现了 92.40% 的平均成功率,在四只未见手上实现了 74.02% 的平均成功率。在真实世界实验中,同一模型在两只未见手上无需额外训练或微调即实现了 76.00% 的总体成功率。这些结果证明了 FunCo-Grasp 在将抓取知识迁移到未见手部方面的有效性。
英文摘要
Cross-embodiment dexterous grasp generation remains challenging because robotic hands differ substantially in geometry, topology, and kinematics. Existing approaches often lack explicit correspondences between structurally different hand regions that play similar functional roles in a grasp, a concept we refer to as functional correspondence. Consequently, their models tend to learn hand-specific interaction patterns rather than transferable grasp knowledge, limiting generalization to unseen hands. To address this limitation, we introduce FunCo-Grasp, which establishes functional correspondences across heterogeneous hand embodiments. Specifically, Functional Part Alignment aligns each hand to a canonical functional schema by mapping physical links to shared functional parts according to their grasping roles, while Canonical Frame Alignment expresses these parts in canonical local frames. These two alignments provide a consistent representation for inter-part and hand-object interactions, allowing the model to learn transferable grasp knowledge across hands. Conditioned on the aligned hand representation and object geometry, a diffusion model generates the target spatial arrangement of the functional parts, which are then converted into an executable joint configuration. Adapting FunCo-Grasp to an unseen hand requires only its geometric and kinematic models and a one-time lightweight functional annotation, without target-hand grasp data, fine-tuning, or learned retargeting. In simulation on held-out objects from the filtered CMapDataset, we achieves average success rates of 92.40% on three seen hands and 74.02% on four unseen hands. In real-world experiments, the same model achieves an overall success rate of 76.00% on two unseen hands without additional training or fine-tuning. These results demonstrate the effectiveness of FunCo-Grasp in transferring grasp knowledge to unseen hands.