arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于视觉语言模型的证据门控任务与运动规划

Evidence-Gated Task and Motion Planning with Vision-Language Models

Tsunehiko Tanaka, Matthew Stephenson, Alistair Macvicar, Edgar Simo-Serra

arXiv 2608.20084首次发表:更新:

发表机构

Waseda University; Flinders University(早稻田大学; 弗林德斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器人执行自然语言指令长时操纵任务时的部分可观测问题,提出 EAFG 框架,通过视觉证据获取与可行性门控提升食谱完成率并减少无效操纵尝试。

AI 中文摘要

执行自然语言指令下长 horizon 操纵任务的机器人需同时推理语义任务结构与几何可行性,但在部分可观测条件下,与目标相关的物体可用性存在不确定性。现有结合视觉语言模型(Vision-Language Models, VLMs)与任务与运动规划(Task and Motion Planning, TAMP)的方法,可能生成依赖 VLM 先验知识却无观测支持的子目标,导致执行失败或意外结果。本文提出证据获取与可行性门控(Evidence Acquisition and Feasibility Gating, EAFG)框架,该框架通过 VLM 生成的探索性子目标与基于 TAMP 的执行获取视觉证据,随后应用可行性门决定是继续任务规划、获取更多证据还是终止。实验表明,在物体使用存在歧义的烹饪任务中,EAFG 通过在规划前发现与任务相关的物体提升了食谱完成率;对于需要不存在物体的指令,EAFG 能做出合适的终止决策,减少对该物体的重复操纵尝试。

英文摘要

Robots executing long-horizon manipulation tasks from natural-language instructions must reason about both semantic task structure and geometric feasibility. However, under partial observability, the availability of goal-relevant objects may be uncertain. In such cases, approaches that combine Vision-Language Models (VLMs) with Task and Motion Planning (TAMP) may generate subgoals that rely on the VLM's prior knowledge without observational support, leading to execution failures or unintended outcomes. We propose Evidence Acquisition and Feasibility Gating (EAFG), a framework that acquires visual evidence through VLM-generated exploratory subgoals and TAMP-based execution. EAFG then applies a feasibility gate to decide whether to proceed with task planning, acquire further evidence, or halt. Our experiments show that, in cooking tasks with ambiguous object use, EAFG improves recipe completion by discovering task-relevant objects before planning. For instructions requiring an absent object, EAFG promotes appropriate halt decisions and reduces repeated attempts to manipulate that object.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑