PDDLCoder:用于LLM辅助符号规划的智能体式PDDL生成工具
PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning
- Technical University of Berlin(柏林工业大学)
- Fraunhofer FOKUS(弗劳恩霍夫FOKUS研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出智能体式PDDL生成框架PDDLCoder,引入含711个问题的NL-pddlgym基准,其在测试集上可生成89.6%可执行规划,优于现有方法。
AI中文摘要:
大型语言模型(LLMs)在长程规划方面仍不可靠,常生成逻辑不一致或不可执行的规划方案。近期的混合方法转而将自然语言翻译为规划领域定义语言(PDDL),使符号规划器能生成可验证的规划方案。但现有方法常依赖僵化的生成流水线、部分PDDL定义或人工反馈,且因缺乏带自动验证的标准化基准,其评估工作受阻。为解决这些局限,我们提出PDDLCoder,一种从自然语言生成PDDL的智能体式框架,可迭代生成、分析并优化规划规范。我们还引入NL-pddlgym,这是一个包含23个领域共711个规划问题的基准数据集,配有可执行的gym环境,用于自动验证规划方案的可执行性。在包含4个保留领域共106个问题的NL-pddlgym测试集上的实验显示,PDDLCoder为89.6%的测试规划问题生成了可执行的规划方案,优于我们对先前PDDL生成方法的适配版本(最高达45.3%),也超过了直接LLM规划方法(在同一测试集上最高达74.5%)。本研究证明了智能体式PDDL生成在规划中的有效性,并为未来LLM辅助符号规划的研究建立了可复现的基准。
英文摘要:
LLMs remain unreliable for long-horizon planning, often generating logically inconsistent or non-applicable plans. Recent hybrid methods instead translate natural language into the Planning Domain Definition Language (PDDL), allowing symbolic planners to produce verifiable plans. However, existing methods frequently rely on rigid generation pipelines, a partial PDDL definition, or human feedback. Furthermore, their evaluation is hindered by the lack of standardized benchmarks with automated verification. To address these limitations, we present PDDLCoder, an agentic framework for PDDL generation from natural language that iteratively generates, analyzes, and refines planning specifications. We further introduce NL-pddlgym, a benchmark dataset comprising 711 planning problems across 23 domains with executable gym environments for the automated verification of plan applicability. Experiments on the NL-pddlgym test set containing 106 problems across 4 held-out domains show that PDDLCoder generates applicable plans for 89.6\% of tested planning problems. This improves upon our adaptations of previous PDDL generation methods, which achieved up to 45.3\%, and outperforms direct LLM planning approaches, which reached up to 74.5\% on the same test set. Our work demonstrates the effectiveness of agentic PDDL generation for planning and establishes a reproducible benchmark for future research on LLM-assisted symbolic planning.