AI 中文总结
本文评估开放权重大语言模型仅用LLM方法修复PDDL模型错误的能力,发现其F1分数优于符号基线,但测试通过率仍不足,无法保证满足修复所需的测试约束。
AI 中文摘要
AI规划旨在寻找一系列能达成指定目标的动作序列,它依赖以规划领域定义语言(PDDL)表示的显式世界模型。当前活跃的研究方向之一是如何检测并修复这类模型中的错误,例如用户提供作为解决方案的正测试计划,以及执行失败的负测试计划,自动化修复方法会修改PDDL模型以满足这些约束。本文评估了近期开放权重大语言模型仅通过LLM方法执行该修复任务的能力。实验表明,符号基线的F1分数为0.49,而表现最佳的LLM在高推理成本下达到0.87,绝对提升0.38;但该设置的平均测试通过率仅为0.82,在Thoughtful领域降至0.06,即使包含测试轨迹的最佳设置也仅达0.92。因此,当前开放权重模型无法保证满足可靠自动化模型修复所需的测试约束。
英文摘要
AI planning is concerned with finding a sequence of actions that achieves a specified goal. It relies on explicit models of the world, commonly represented in the Planning Domain Definition Language (PDDL). An active line of research investigates how errors in such models can be detected and repaired. For example, users may provide positive test plans that are solutions, and negative test plans that fail during execution. Automated repair methods then modify the PDDL model to satisfy these constraints. In this paper, we evaluate the ability of recent open-weight large language models to perform this repair task using an LLM-only approach. Our experiments show that the symbolic baseline achieves an $F_1$ score of $.49$, while the best-performing LLM reaches $.87$ with high reasoning effort, an absolute improvement of $.38$. However, that setting has a mean test pass rate of only $.82$, falling to $.06$ on the Thoughtful domain; even the best setting that includes the test traces reaches only $.92$. Thus, current open-weight models cannot guarantee satisfaction of the test constraints required for reliable automated model repair.