发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种结合语言任务上下文与视觉接地模仿学习的框架,解决机器人拆解中多任务长时程执行问题,无需显式标注,端到端成功率较基线提升35至75个百分点。
AI 中文摘要
现实世界中的机器人拆解需要长时间跨度的执行,机器人必须在单个场景中对多个部件执行有序的操作任务序列。多个有效的任务目标和多样的装配配置使得模仿策略难以仅从原始观测中推断出预期技能,尤其是当训练数据无法覆盖现实世界配置和部件几何形状的组合多样性时。我们表明,通过语言引入任务上下文可以缓解这些挑战,它为技能选择提供显式结构,并将语言指定的任务与视觉场景中对应的操作目标关联起来。所提出的框架将分层任务选择与任务上下文感知的模仿学习相结合,将语言指令接地到空间视觉表示中,用于机器人拆解。该框架无需显式对象标注即可在多样的连接器几何形状和装配配置中泛化。我们的方法在端到端任务成功率上比基线扩散策略提高了35个百分点,比之前的任务上下文感知基线提高了75个百分点。
英文摘要
Real-world robotic disassembly requires long-horizon execution, where robots must perform ordered sequences of manipulation tasks across multiple parts within a single scene. Multiple valid task goals and diverse assembly configurations make it difficult for imitation policies to infer the intended skill from raw observations alone, particularly when training data cannot cover the combinatorial diversity of real-world configurations and part geometries. We show that incorporating task context through language alleviates these challenges by providing explicit structure for skill selection and associating language-specified tasks with their corresponding manipulation targets in the visual scene. The proposed framework combines hierarchical task selection with task-context-aware imitation learning to ground language instructions in spatial visual representations for robotic disassembly. The resulting framework generalizes across diverse connector geometries and assembly configurations without requiring explicit object annotations. Our method improves end-to-end task success by 35 percentage points over the baseline diffusion policy and by 75 percentage points over the previous task-context-aware baseline.
CommentsAccepted for publication in IEEE Robotics and Automation Letters (RA-L), 2026. 8 figures