arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CosmoH2G:面向复杂空间运动物体操作的手到夹爪迁移数据集与基线方法

CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements

Hongxiang Zhao, Mutian Xu, Zeyu Jin, Yiming Hao, Shuguang Cui, Xiaoguang Han

arXiv 2609.07498首次发表:更新:

发表机构

SSE, CUHKSZ; GenuX; FNii-Shenzhen(香港中文大学(深圳)理工学院; 真象科技; 鹏城国家实验室(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有手到夹爪迁移方法难以处理复杂空间运动的问题,提出包含大规模配对数据集和两阶段关键帧预测框架的CosmoH2G方法,在仿真与真实实验中显著优于传统基线。

AI 中文摘要

将人类手部演示迁移到机器人夹爪近期已成为机器人学习的一种经济高效的解决方案。然而,现有方法大多局限于简单、平面任务,无法处理机器人操作所必需的复杂空间运动(例如涉及旋转或翻转的复杂轨迹)。受这一差距的驱动,我们采用了一种由细粒度手部姿态运动引导的隐式、数据驱动方法。为此,我们引入了一个可扩展的数据采集流程,用于收集手-夹爪配对演示,该流程遵循一项优先考虑运动复杂性的严格协议,并利用手持式夹爪实现无缝动作模仿。这产生了一个大规模配对数据集,包含1,254个独特物体上的6,189个片段,其空间复杂性显著高于现有基准。然而,学习这种复杂映射仍然具有挑战性。我们观察到,朴素的全夹爪姿态序列端到端生成是不够的,因为微小的轨迹偏差在复杂动力学下会迅速累积。为解决此问题,我们提出了一个两阶段框架:第一阶段预测稀疏的夹爪关键帧(初始和终止)以简化映射目标,而第二阶段基于这些关键帧生成完整的连续动作序列。此外,为减轻累积漂移,我们保持夹爪的朝向被学习,同时基于抓取启发式和运动学一致性对其平移进行后优化。在仿真和真实机器人实验中,我们的框架实现了复杂空间操作的稳定且精确的手到夹爪迁移,显著优于传统基线。项目页面:此 https URL。

英文摘要

Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.

CommentsSIGGRAPH Aisa 2026; Project page: https://cosmoh2g.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑