arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向自然语言到PDDL问题生成的接地评估与修复

Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

Joana Rosa, Pedro Santos, Valdemar Oliveira, Romão Silva, L. Miguel Silveira, Bruno Martins

arXiv 2609.09898首次发表:更新:

发表机构

INESC INOV; INESC ID; Instituto Superior Técnico, Universidade de Lisboa; Motamineral Minerais Industriais S.A.(INESC INOV; INESC ID; 里斯本大学高等技术学院; Motamineral 工业矿产公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究NL到PDDL生成流水线,结合LLM生成与多种检查及迭代修复,发现操作成功与基准参考重建存在显著分歧,结构化修复有效但PDDL 2.1参考重建仍具挑战。

AI 中文摘要

大型语言模型(LLMs)在将自然语言(NL)规划描述转换为PDDL问题实例方面展现出潜力。然而,诸如语法有效性或规划器成功等标准评估标准可能显著高估对所述任务的忠实度:生成的问题可能可解析且可求解,同时却错误地表示了预期的初始状态、目标、对象结构或优化目标。本文研究了一个端到端的NL-to-PDDL流水线,该流水线结合了LLM生成、基于PDDL解析、规划和验证的检查、领域一致性检查器、LLM评论器以及迭代修复。细粒度的修复反馈由领域描述、生成的问题、自然语言问题描述和操作诊断构建。基于参考的对比用于对策划的基准PDDL问题描述进行事后基准分析,这些离线检查包括重命名不变的结构匹配和语义等价性(在领域支持可用的情况下)。在Planetarium、AutoPlanBench和策划的PDDL 2.1问题上,结果表明操作成功与基准参考重建可能显著分歧。结果还表明,结构化修复可能有用,并且PDDL 2.1对于参考重建仍然具有挑战性,即使操作成功有所改善。

英文摘要

Large Language Models (LLMs) have shown promise for translating Natural Language (NL) planning descriptions into PDDL problem instances. However, standard evaluation criteria such as syntactic validity or planner success can substantially overestimate faithfulness to the described task: a generated problem may be parseable and solvable while misrepresenting the intended initial state, goal, object structure, or optimization target. This paper studies an end-to-end NL-to-PDDL pipeline that combines LLM generation, checks in terms of PDDL parsing, planning and validation, a domain-conformance checker, an LLM critic, and iterative repair. Fine-grained repair feedback is constructed from the domain description, the generated problem, the natural language problem description, and operational diagnostics. Reference-based comparisons against curated benchmark PDDL problem descriptions are used for post-hoc benchmark analysis, and these offline checks include renaming-invariant structural matching and semantic equivalence, where domain support is available. Across Planetarium, AutoPlanBench, and curated PDDL~2.1 problems, results show that operational success and benchmark-reference reconstruction can diverge substantially. Results also show that structured repair can be useful, and that PDDL~2.1 remains challenging for reference reconstruction, even when operational success improves.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑