arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用质量多样性优化发现多模态具身智能体的多样化规划策略

Discovering Diverse Planning Policies for Multimodal Embodied Agents with Quality-Diversity Optimization

Pengfei Xu, Yong Liu, Xiaoya Nan, Qiang Yang, Peilan Xu

arXiv 2608.08523首次发表:更新:

发表机构

School of Artificial Intelligence, Nanjing University of Information Science and Technology(南京信息工程大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态具身智能体现有规划器易因单一规划风格失效导致停滞的问题,提出质量多样性框架,通过进化策略模板构建多样化策略库,实现自适应规划与故障恢复,在ThreeDWorld基准上提升了任务成功率与交互效率。

AI 中文摘要

多模态具身智能体日益需要通过将视觉观测、文本目标和交互历史整合到闭环决策中来解决长程任务。然而,最先进的基于大模型的规划器在执行过程中往往依赖单一主导规划风格,一旦该执行模式失效,智能体可能会停滞多步,反复与环境交互却无实质进展。为解决此局限,本文提出一种质量多样性(QD)框架,用于发现多模态具身智能体的多样化规划策略。该方法将规划策略模板视为可进化个体,组织成按行为索引的存档,而非将搜索压缩为单一提示风格。在离线阶段,将回滚轨迹总结为结构化的成功与失败经验,通过重组和经验引导的变异来指导策略变异;生成的策略被映射到由交互强度和目标导向性定义的行为空间,每个生态位中质量最高的策略被保留在存档中。在线阶段,智能体一次执行一个策略,同时监控任务进度;当检测到持续停滞时,系统回滚到最新检查点,切换到存档中行为不同的策略以恢复执行。在ThreeDWorld运输基准上的实验表明,与代表性基线规划器相比,所提框架提升了任务成功率和交互效率。这些结果表明,发现多样化策略库是支持自适应多模态规划和在线故障恢复的有效方式。

英文摘要

Multimodal embodied agents are increasingly required to solve long-horizon tasks by integrating visual observations, textual goals, and interaction history into closed-loop decision making. However, state-of-the-art large-model-based planners often rely on a single dominant planning style during execution. Once this execution mode becomes ineffective, the agent may remain stalled for many steps, repeatedly interacting with the environment without making meaningful progress. We address this limitation by proposing a Quality-Diversity (QD) framework for discovering diverse planning policies for multimodal embodied agents. The proposed method treats planning-policy templates as evolvable individuals and organizes them into a behavior-indexed archive rather than collapsing search to a single prompt style. In the offline stage, rollout trajectories are summarized into structured success and failure experiences, which guide policy variation through recombination and experience-guided mutation. The resulting policies are mapped into a behavior space defined by interaction intensity and goal-directedness, and the highest-quality policy in each niche is retained in the archive. In the online stage, the agent executes one policy at a time while monitoring task progress. When persistent stall is detected, the system rolls back to the latest checkpoint and switches to a behaviorally distinct archive policy to resume execution. Experiments on the ThreeDWorld transport benchmark show that the proposed framework improves both task success and interaction efficiency over representative baseline planners. These results suggest that discovering diverse policy repertoires is an effective way to support adaptive multimodal planning and online failure recovery.

CommentsTo appear in ACM MM (MM '26)

DOI:10.1145/3767308.3836597

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑