通过状态门控专家弥合手持与遥操作监督在接触丰富操作中的差距
Bridging Handheld and Teleoperated Supervision for Contact-Rich Manipulation via State-Gated Experts
查看机构详情
- RAI Institute(RAI研究所)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对接触丰富操作任务,提出BRIDGE方法,通过状态门控混合扩散策略专家,融合手持数据与少量遥操作演示,在接触敏感阶段利用期望动作,提升成功率最高36.7%。
中文摘要 AI 辅助
手持数据采集系统(如通用操作接口UMI)能够在多样环境中实现可扩展的数据采集,但仅捕获观察到的动作而非机器人控制器执行的期望动作。相比之下,遥操作直接捕获期望动作,但采集过程耗时过多。我们通过任务阶段中动作有效性的视角重新审视这一权衡。我们观察到,手持轨迹在容忍的自由空间阶段提供有效监督,但在接触敏感阶段缺乏动态可行性,因为以高刚度跟踪观察到的轨迹会产生大而不安全的接触力。我们研究了这两种监督类型在接触丰富操作中的相互作用,发现结合手持数据与少量针对性遥操作演示的训练策略提供了一种高效的混合策略。具体而言,我们并非对整个任务进行遥操作,而是仅对基础手持策略失败的任务片段收集部分遥操作演示。然而,简单混合手持和遥操作阶段特定数据会导致性能低于仅使用手持数据训练。为解决观察监督与期望监督之间的不匹配,我们提出了通过门控专家进行模仿数据的双模态路由(BRIDGE),这是一种混合扩散策略专家,根据当前机器人状态路由到专门的阶段任务头。值得注意的是,我们的方法能够在接触敏感阶段实现任务阶段特定的期望动作使用,并在三个接触丰富操作任务中,相较于仅手持基线将成功率提升高达36.7%。
英文摘要
Handheld data collection systems, such as the Universal Manipulation Interface (UMI), enable scalable data collection across diverse environments but only capture observed actions rather than the desired actions executed by a robot controller. In contrast, teleoperation captures desired actions directly, but requires a robot in-the-loop. We revisit this trade-off through the lens of action validity across task phases. We observe that handheld trajectories provide valid supervision in tolerant, free-space phases, but can be poor supervision targets for contact-sensitive phases, requiring high stiffness to minimize tracking errors, resulting in large contact forces. We study the interaction between these two supervision types for contact-rich manipulation and find that training policies that combine handheld data with a small number of targeted teleoperated demonstrations provide an effective hybrid strategy. Specifically, rather than teleoperating the entire task, we only collect partial teleoperated demonstrations for task segments where base handheld policies fail. However, naively mixing handheld and teleoperated phase-specific data yields worse performance than training on handheld data alone. To address this mismatch between observed and desired supervision, we propose Bi-modal Routing for Imitation Data via Gated Experts (BRIDGE), a mixture of diffusion policy experts that routes between specialist task phase heads conditioned on the current robot state. Notably, our approach enables task-phase specific use of desired actions during contact sensitive segments and improves success rates over handheld-only baselines by up to 36.7% across three contact-rich manipulation tasks.