Planetarium:用于将文本转换为结构化规划语言的严格基准
Planetarium: A Rigorous Benchmark for Translating Text to Structured Planning Languages
- Brown University(布朗大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有文本转结构化规划语言评估方法的不足,本文提出Planetarium基准,通过新颖PDDL等价算法和大规模数据集评估语言模型生成PDDL的能力,揭示任务复杂性并凸显严格基准的必要性。
AI中文摘要:
近期研究探索了语言模型在规划问题中的应用。其中一种方法是将规划任务的自然语言描述转换为结构化规划语言,例如规划领域定义语言(PDDL)。现有评估方法难以确保语义正确性,且依赖简单或不现实的数据集。为填补这一空白,我们引入Planetarium基准,旨在评估语言模型从规划任务的自然语言描述生成PDDL代码的能力。Planetarium包含一种新颖的PDDL等价算法,可灵活评估生成PDDL的正确性,以及一个涵盖73种独特状态组合、难度各异的145,918个文本到PDDL对的数据集。最后,我们评估了多个API访问型和开源语言模型,揭示了该任务的复杂性。例如,GPT-4o生成的PDDL问题描述中,96.1%语法可解析,94.4%可求解,但仅24.8%语义正确,凸显了为该问题建立更严格基准的必要性。
英文摘要:
Recent works have explored using language models for planning problems. One approach examines translating natural language descriptions of planning tasks into structured planning languages, such as the planning domain definition language (PDDL). Existing evaluation methods struggle to ensure semantic correctness and rely on simple or unrealistic datasets. To bridge this gap, we introduce \textit{Planetarium}, a benchmark designed to evaluate language models' ability to generate PDDL code from natural language descriptions of planning tasks. \textit{Planetarium} features a novel PDDL equivalence algorithm that flexibly evaluates the correctness of generated PDDL, along with a dataset of 145,918 text-to-PDDL pairs across 73 unique state combinations with varying levels of difficulty. Finally, we evaluate several API-access and open-weight language models that reveal this task's complexity. For example, 96.1\% of the PDDL problem descriptions generated by GPT-4o are syntactically parseable, 94.4\% are solvable, but only 24.8\% are semantically correct, highlighting the need for a more rigorous benchmark for this problem.