发表机构
Humanoid Robots Lab and Center for Robotics, University of Bonn; Lamarr Institute for Machine Learning and Artificial Intelligence(人形机器人实验室及机器人技术中心,波恩大学; 拉马尔机器学习与人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究长期机器人重排任务,提出基于未标记子任务演示的隐式行为协调方法,通过价值引导动作选择学习类技能行为,实验表明该方法在复杂任务中表现优,能处理更大行为库,为显式技能管道提供替代。
AI 中文摘要
长期的机器人重排任务常被视为技能排序问题,需预定义技能、技能标签或边界以及特定任务的切换逻辑。随着行为数量和任务范围增加,这种显式技能抽象难以扩展。本文将重排任务表述为基于未标记子任务演示的隐式行为协调,从混合行为数据中直接学习类技能行为,并通过价值引导的动作选择进行协调。在Habitat重排任务中的实验从三方面支持了该方法:在更复杂任务中优于特定任务模仿基线;消融实验表明可靠的评论家引导候选选择对协调多模态行为至关重要;扩展实验表明该方法能处理更大行为库且性能更强。结果表明显式技能抽象并非长期重排的先决条件,隐式行为协调为基于显式技能的管道提供了有前景的数据驱动替代方案。
英文摘要
Long-horizon robotic rearrangement is commonly formulated as a skill-sequencing problem, where distinct behaviors are explicitly represented and coordinated by a planner or high-level policy. We investigate whether such explicit behavior identities and sequencing interfaces are necessary at all. We introduce implicit behavior coordination from sub-task demonstrations, where separately collected behaviors are coordinated without behavior identity labels, complete-task demonstrations, or task-ordering supervision. Our key observation is that overlap between sub-task demonstrations induces multimodal action distributions that need not be resolved through explicit behavior partitioning. Instead, this overlap-induced multimodality can be exploited as a coordination resource. We instantiate this idea with a shared Flow Matching policy that preserves multiple action modes and critic-guided in-sample planning that propagates task value across demonstrations and selects task-relevant modes. Experiments in Habitat and on a real robot show that implicit behavior coordination remains effective under reduced cross-behavior overlap, larger behavior mixtures, longer horizons, and execution failures, supporting the idea that long-horizon coordination can emerge directly from sub-task demonstrations without explicitly recovering or sequencing behavior identities.