可指令智能体的规划与控制解耦
Decoupling Planning and Control for Instructable Agents
浏览论文内容
中文总结 AI 辅助
本研究提出Instruct-to-Act系统,将VLM规划器与世界模型控制器解耦,在七个具身环境中验证其性能优于仅控制器、直接VLM动作生成等基线,可替换VLM规划器且保持快速控制。
中文摘要 AI 辅助
近期研究显示,经预训练、指令微调的视觉语言模型(VLM)能较好地将指令与观测映射为高层规划,但难以在陌生环境中将此类规划转化为可靠、低延迟的动作序列。与此同时,世界模型控制器擅长快速的观测到动作的控制,但缺乏开放式任务引导。本研究将这些优势结合为单一系统Instruct-to-Act,我们训练世界模型控制器,使其在由VLM规划器生成的稀疏、高延迟、高层文本指令条件下自主高频行动。为训练控制器具备指令可执行性,我们用合成指令重新标记控制器策略回滚的片段,并联合优化行为克隆目标,以及现有的奖励最大化和世界建模目标。我们在七个具身环境中评估所提方法,包括三个多智能体环境,其中VLM规划器通过语言协调,而训练后的控制器作为其执行器。在匹配的观测与动作空间下,我们的解耦方法始终优于仅控制器和直接VLM动作生成的变体,保持快速控制,且无需微调即可替换不同预训练VLM规划器,在七项任务中的六项上与强大的视觉语言动作及多智能体强化学习基线具有竞争力。
英文摘要
Recent work shows that pre-trained, instruction-tuned vision-language models (VLMs) perform well at mapping from instructions and observations to high-level plans, but struggle to realize such plans as reliable low-latency action sequences in unfamiliar environments. At the same time, world-model controllers excel at fast observation-to-action control, but lack open-ended task guidance. In this work, we combine these strengths into a single system, Instruct-to-Act, where we train a world-model controller to act autonomously at high frequency when conditioned on sparse, higher-latency, and high-level text instructions generated by a VLM planner. To train controllers to be language-instructable, we relabel segments of controller policy rollouts with synthetic instructions and jointly optimize a behavior-cloning objective along with existing reward-maximizing and world-modeling objectives. We evaluate our proposed approach across seven embodied environments, including three multi-agent environments where VLM planners coordinate through language while trained controllers serve as their actuators. Under matched observation and action spaces, our decoupled approach consistently outperforms controller-only and direct VLM action-generation variants, preserves fast control, and lets us swap in different pretrained VLM planners without fine-tuning, while remaining competitive with strong vision-language-action and multi-agent RL baselines on six of seven tasks.
发表机构
- UC Berkeley(加州大学伯克利分校)
- UBC(不列颠哥伦比亚大学)
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。