arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过交互中心建模实现双指夹爪操作的统一跨域表示

Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling

Guanlin Li, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li, Chen Zhao, Tong Yang, Jie Tang, Jing Zhang

arXiv 2609.31207首次发表:更新:

发表机构

Renmin University of China; Zhipu AI; Tsinghua University(中国人民大学; 智谱AI; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出交互中心建模框架,通过通用夹爪抽象和混合特征实现双指夹爪操作的跨域表示,首次在模仿学习中同时达到竞争性基准分数和跨具身零样本迁移。

AI 中文摘要

在模仿学习中实现鲁棒的跨具身泛化,需要克服一个关键的表示缺陷,该缺陷将任务语义与硬件特定的视觉几何不可避免地纠缠在一起。我们提出了一种以交互为中心的框架,通过参数化的通用夹爪抽象利用双指夹爪的共享结构,从而产生一个规范的夹爪框架表示。给定语言和RGB-D观测,VLM推断子任务并确定交互三元组(夹爪、持有物、目标),而SAM 2.1跟踪掩码以减少VLM查询。我们设计了简洁的混合特征,将目标/碰撞人工势场用于全局引导,与分割的夹爪框架点云用于局部几何相结合,并使用Flow-Matching Transformer预测平滑的7自由度动作块。在仿真和真实世界任务中的实验表明,我们是第一个同时实现竞争性基准分数和极端跨具身/跨视角零样本模拟到现实迁移到完全不同的异构机器人平台的模仿学习方法。

英文摘要

Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a parameterized universal gripper abstraction, yielding a canonical gripper-frame representation. Given language and RGB-D observations, a VLM infers the subtask and grounds an interaction triplet (gripper, held, target), while SAM~2.1 tracks masks to reduce VLM queries. We design concise hybrid features that combine target/collision artificial potential fields for global guidance with segmented gripper-frame point clouds for local geometry, and use a Flow-Matching Transformer to predict smooth 7-DoF action chunks. Experiments in simulation and real-world tasks demonstrate that ours is the first imitation learning approach to simultaneously achieve competitive benchmark scores and extreme cross-embodiment/cross-viewpoint zero-shot sim-to-real transfer to completely distinct, heterogeneous robot platforms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑