动态规划:部分可观测操纵任务中用于TAMP执行的事件触发式基础模型规划
Plan Along the Way: Event-Triggered Foundation-Model Planning for TAMP Execution in Partially Observable Manipulation
浏览论文内容
中文总结 AI 辅助
本文提出ROBUST TAMP框架,针对部分可观测操纵任务,通过事件触发式重规划机制提升TAMP执行的任务成功率,验证了其在6个RLBench/CoppeliaSim场景中的有效性。
中文摘要 AI 辅助
部分可观测环境中的操纵任务需要在场景信息不完整的情况下进行规划。在此类场景中,初始有效的计划可能执行成功,但仍不足以完成任务。现有的基础模型引导的任务与运动规划(TAMP)系统可生成有用的长程任务分解、子目标或约束,但它们通常假设可获取完全指定的场景状态,或在子目标、细化操作或执行尝试失败后调用模型级重规划。本文提出ROBUST TAMP,这是一种模块化的大语言模型(LLM)/视觉语言模型(VLM)引导的反应式TAMP规划框架,适用于执行过程中可能出现未被观测到的任务相关对象及非目标对象的场景。该框架将基础模型规划器限制在当前可见的关系型场景状态范围内,针对严格的可执行接口验证生成的任务级动作,并将已接受的动作路由至特定场景的执行适配器。对象发现被视为独立的重规划事件,在稳定的执行周期后,系统会利用已完成动作的历史记录和结构化重规划事件上下文,重建可见场景状态并进行重规划。评估在6个RLBench/CoppeliaSim厨房和烧烤变体场景中开展,这些场景涉及隐藏对象、非目标对象发现、关节容器交互及时序操纵流程。我们在相同的验证、执行、监控和重规划流程下,对比了不同规模的纯文本LLM规划器与VLM规划器,报告了任务成功率、部分目标完成度、由发现和失败触发的重规划行为、隐式非目标对象处理情况及规划器推理成本。
英文摘要
Manipulation in partially observable environments requires planning under incomplete scene information. In such settings, an initially valid plan may execute successfully yet remain insufficient for task completion. Existing foundation-model-guided task and motion planning (TAMP) systems can generate useful long-horizon task decompositions, subgoals, or constraints, but they often assume having access to a fully specified scene state or invoke model-level replanning after a subgoal, refinement, or execution attempt fails. We present ROBUST TAMP, a modular LLM/VLM-guided planning framework for reactive TAMP where unseen task-relevant and non-target objects may become visible during execution. The framework restricts the foundation-model planner to the currently visible relational scene state, validates generated task-level actions against a strict executable interface, and routes the accepted actions to scene-specific execution adapters. Object discovery is treated as a distinct replanning event and, after a stable execution horizon, the system reconstructs the visible scene state and replans using completed-action history and structured replanning event context. Evaluations are performed on six RLBench/CoppeliaSim kitchen and grill variants involving hidden objects, non-target object discovery, articulated-container interaction, and temporal manipulation procedures. We compare text-only LLM and VLM planners of different sizes under the same validation, execution, monitoring, and replanning pipeline, reporting task success, partial goal completion, discovery- and failure-triggered replanning behavior, implicit non-target-object handling, and planner inference cost.
发表机构
- Robotics Research Center, IIIT Hyderabad(IIIT海得拉巴机器人研究中心)
- Center for Security, Theory and Algorithmic Research, IIIT Hyderabad(IIIT海得拉巴安全、理论与算法研究中心)
机构由 AI 辅助整理,请以论文原文为准。