arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

哪个LLM最适合将自然语言目标翻译为PDDL

Which LLM is Best for Translating Natural Language Goals to PDDL

Tomas Balyo, Lukas Chrpa, G. Michael Youngblood

arXiv 2609.18731首次发表:更新:

发表机构

Filuta AI, Inc.(Filuta AI公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文评估六个LLM将自然语言目标翻译为PDDL的能力,提出提示模板,发现所有模型正确率超92%,Gemini 2.5 Flash准确率最高,GPT-4.1速度最快,但仍存在语言歧义等挑战。

AI 中文摘要

在自动规划中,弥合人类意图与机器执行之间的差距仍然是一个挑战,因为用PDDL等正式语言表达目标限制了非专家的可访问性。本文通过实验评估了当前大型语言模型(LLMs)是否能够可靠地将视频游戏测试人员用非正式语言编写的自然语言测试目标,翻译成适合经典规划的格式良好的PDDL目标。我们提出了一个精心设计的提示模板,整合了迭代实验中的见解,旨在最大化多个最先进LLMs的准确性和响应连贯性。我们使用真实世界、特定领域的基准,系统评估了六个当代模型在正确性、速度和错误倾向方面的表现。所有模型均表现出高正确性,超过92%,其中Gemini 2.5 Flash以96%的准确率位居最高,且误报发生率最低,而GPT-4.1在响应速度上领先。尽管取得了这些进展,模型性能仍存在关键差异,偶尔的失败源于语言歧义和领域表示的局限性。我们的分析强调了LLMs在作为自然语言目标与自动规划管道之间稳健桥梁方面的显著进展和持续存在的差距。

英文摘要

Bridging the gap between human intent and machine execution remains a challenge in automated planning, where expressing goals in formal languages like PDDL restricts accessibility to non-experts. This paper empirically evaluates whether current Large Language Models (LLMs) can reliably translate natural language testing goals, written in informal language by video game testers, into well-formed PDDL targets suitable for classical planning. We present a carefully designed prompt template, integrating insights from iterative experimentation, aimed at maximizing both accuracy and response coherence from multiple state-of-the-art LLMs. Six contemporary models are systematically assessed on correctness, speed, and error tendencies using real-world, domain-specific benchmarks. All models demonstrate high correctness, exceeding 92\%, with Gemini 2.5 Flash achieving the highest accuracy at 96\% and the lowest incidence of false positives, while GPT-4.1 leads in response speed. Despite these advances, critical distinctions exist in model performance, and occasional failures arise from language ambiguity and limitations in domain representation. Our analysis underscores both the significant progress and ongoing gaps in enabling LLMs to act as robust bridges between natural language objectives and automated planning pipelines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑