发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PDDL-ART是一种基于VLM的框架,可自主从演示生成PDDL描述,经多阶段校正确保语义对齐,在发动机维护和家庭操纵任务上平均成功率达93.3%,优于基线规划器。
AI 中文摘要
使用PDDL的符号规划为长时程机器人操纵提供了原则性框架,但构建准确的PDDL领域和问题描述仍是重大瓶颈,通常需要大量领域专业知识。我们提出一种基于视觉语言模型(VLM)的方法,名为PDDL-ART,该框架可从单个专家演示、自然语言任务描述以及可用高级动作名称库中自主生成特定任务的PDDL领域和问题描述。PDDL-ART不需要任何领域模板、动作签名或微调。为确保生成的描述不仅语法有效,还与演示任务语义对齐,PDDL-ART引入了在语法、语义和执行层面运行的多阶段校正流水线。执行引导校正的关键组件是符号谓词接地。PDDL-ART不仅依赖视觉观测,而是利用现代VLM的工具使用能力,结合几何和时间推理来评估无法仅从图像中直接辨别的关系谓词。关键在于,该模型自主决定何时调用这些工具以及如何解释其输出。我们在发动机维护和家庭领域的具有挑战性的操纵任务上评估PDDL-ART,包括需要记忆、抽象谓词推理以及目标状态与初始状态视觉上无法区分的任务。PDDL-ART的平均成功率达到93.3%,而基于VLM的基线规划器的平均成功率为78.3%。
英文摘要
Symbolic planning with PDDL offers a principled framework for long-horizon robot manipulation, but constructing accurate PDDL domain and problem descriptions remains a significant bottleneck, typically requiring substantial domain expertise. We present a Vision-Language Model (VLM)-based approach called PDDL-ART, a framework that autonomously generates task-specific PDDL domain and problem descriptions from a single expert demonstration, a natural language task description, and a library of available high-level action names. PDDL-ART does not require any domain templates, action signatures, or fine-tuning. To ensure the generated descriptions are not only syntactically valid but semantically aligned with the demonstrated task, PDDL-ART introduces a multi-stage correction pipeline operating at syntactic, semantic, and execution levels. A key component of execution-guided correction is symbolic predicate grounding. Instead of relying solely on visual observations, PDDL-ART leverages the tool-use capabilities of modern VLMs to incorporate geometric and temporal reasoning for evaluating relational predicates that are not directly discernible from images alone. Critically, the model autonomously determines when to invoke these tools and how to interpret their outputs. We evaluate PDDL-ART on challenging manipulation tasks in engine maintenance and household domains, including tasks that require memory, abstract predicate inference, and goal states that are visually indistinguishable from the initial state. PDDL-ART achieves an average success rate of 93.3%, compared to 78.3% for a baseline VLM-based planner.