从失败到监督:用于鲁棒长视距具身规划的DynamicEnvPlan
From Failures to Supervision: DynamicEnvPlan for Robust Long-Horizon Embodied Planning
浏览论文内容
中文总结 AI 辅助
该研究提出DynamicEnvPlan闭环框架,通过规划、扰动与防护校正模块生成恢复轨迹并分阶段微调,在104个任务-场景组合上使具身规划成功率从33.3%升至76.2%,提升了物理世界交互的鲁棒性。
中文摘要 AI 辅助
物理世界交互本质上是动态的,环境在执行过程中会发生演化,要求智能体在非平稳条件下调整其规划。我们研究环境偏差与执行不确定性下的长视距具身规划这一挑战。现有具身任务基准可暴露此类失败,但这些失败通常被视为评估结果,而非训练智能体恢复能力的可学习信号。本研究提出DynamicEnvPlan,一种用于动态环境中高层规划的闭环框架,它将具身任务执行与人形智能体、高层原语技能、结构化语义记忆及可控扰动相结合。我们的数据合成设计包含规划、扰动与防护校正模块,可将动态执行状态转化为面向恢复的轨迹,所得轨迹用于分阶段监督微调,使规划器能从正常执行与扰动恢复轨迹中学习。我们使用涵盖独立同分布、组合泛化与分布外设置的104个任务-场景组合进行微调与评估,结果显示DynamicEnvPlan将基础规划器的成功率从33.3%提升至76.2%,同时在物理世界交互的所有7项关键评估指标(包括安全性与可供性合规性)上均有改善。
英文摘要
Physical-world interaction is inherently dynamic, as environments can evolve during execution, requiring agents to adapt their plans under non-stationary conditions. We study this challenge through long-horizon embodied planning under environment deviations and execution uncertainty. Existing embodied-task benchmarks can expose such failures, but these failures are usually treated as evaluation outcomes instead of learnable signals for training agents to recover. In this work, we introduce DynamicEnvPlan, a closed-loop framework for high-level planning in dynamic environments. It extends embodied task execution with humanoid agents, high-level primitive skills, structured semantic memory, and controllable perturbations. Our data synthesis design consists of planning, perturbation, and guarded correction modules that turn dynamic execution states into recovery-oriented traces. The resulting traces are used for staged supervised fine-tuning, enabling the planner to learn from both nominal execution and perturbed recovery trajectories. Using 104 task-scene combinations spanning i.i.d., compositional generalization, and out-of-distribution settings for fine-tuning and evaluation, DynamicEnvPlan boosts success rate from 33.3% for the base planner to 76.2%, while improving across all seven evaluation metrics critical to physical-world interaction, including safety and affordance compliance.
发表机构
- University of Chinese Academy of Sciences(中国科学院大学)
- Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。