发表机构
Robotics Institute, Carnegie Mellon University; Keio University; College of Engineering, University of California, Berkeley; Bosch Center for Artificial Intelligence(卡内基梅隆大学机器人学院; 庆应义塾大学; 加州大学伯克利分校工程学院; 博世人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SkillWeave是一种异构演示框架,结合遥操作与运动觉示教,通过物体掩码条件扩散策略和后继感知终端引导,在三个真实世界长视距灵巧操作任务中取得了27%的平均端到端成功率。
AI 中文摘要
灵巧操作既需要大规模的任务推进,又需要精确的接触式交互,这使得收集能同时支持这两种场景的演示极具挑战性。我们提出SkillWeave,一种用于长视距灵巧操作的异构演示框架,它结合了遥操作(用于粗略的到达和搬运)与运动觉示教(用于精确的、接触式的技能)。为解决运动觉数据收集过程中演示者存在所引入的视觉不匹配问题,我们提出了一种物体掩码条件扩散策略,该策略使用离线物体分割作为训练监督,并在部署时使用轻量级学习掩码预测器,避免了在线分割和图像修复。为缓解独立训练的子任务策略之间的分布偏移,我们引入了后继感知终端引导,它从先前策略采样的动作中进行选择,以引导系统到达后继者演示的初始状态分布所支持的状态。在三个真实世界长视距任务中,SkillWeave实现了27%的平均端到端成功率;掩码条件运动觉策略将灵巧子任务成功率提升至平均65%,而后继感知交接实现了平均87%的组合效率。这些结果表明,使演示模态与交互场景匹配、明确解决运动觉视觉不匹配问题、以及将策略交接引导至后继者支持的状态,可显著提升长视距灵巧操作性能。视频和代码可在该http链接获取。
英文摘要
Dexterous manipulation requires both large-scale task progression and precise contact-rich interaction, making it challenging to collect demonstrations that effectively support both regimes. We present SkillWeave, a heterogeneous demonstration framework for long-horizon dexterous manipulation that combines teleoperation for coarse reaching and transport with kinesthetic teaching for precise, contact-rich skills. To address the visual mismatch introduced by the demonstrator's presence during kinesthetic data collection, we propose an object-mask-conditioned diffusion policy that uses offline object segmentation for training supervision and a lightweight learned mask predictor at deployment, avoiding online segmentation and image inpainting. To mitigate distribution shift between independently trained sub-task policies, we introduce successor-aware terminal steering, which selects among actions sampled from the predecessor policy to guide the system toward states supported by the successor's demonstrated initial-state distribution. Across three real-world long-horizon tasks, SkillWeave achieves 27% average end-to-end success. Mask-conditioned kinesthetic policies improve dexterous sub-task success to an average of 65%, while successor-aware handoffs achieve an average composition efficiency of 87%. These results show that matching demonstration modality to interaction regime, explicitly addressing kinesthetic visual mismatch, and steering policy handoffs toward successor-supported states substantially improves long-horizon dexterous manipulation. Videos and code are available at skillweave-authors.github.io .
CommentsIROS 2026 Workshop on Haptics and Dexterous Manipulation