arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当我们谈论大语言模型规划时我们在谈论什么:两种不同规划能力的证据

What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities

Sukai Huang, Chenyuan Zhang, Fucai Ke, Zhixi Cai, Naim Rastgoo, Gholamreza Haffari, Hamid Rezatofighi

arXiv 2607.11197首次发表:更新:

发表机构

Faculty of Information Technology, Monash University(莫纳什大学信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型规划任务性能差异原因,通过在ACPBench-Hard上评估多个模型家族,应用多维项目反应理论模型,发现操作推理和结构枚举两个塑造规划性能的维度,推动大语言模型规划的能力层面评估。

AI 中文摘要

当大语言模型在规划任务中表现出参差不齐的性能时,这些差距通常归因于任务难度。我们认为这种解释是不完整的,因为任务层面的差异可能反映了不同的潜在规划能力,而不是单一能力范围内的差异。我们在ACPBench-Hard上研究这个问题,通过在不同的测试时推理预算下评估多个大语言模型家族,并应用多维项目反应理论模型来揭示大语言模型规划背后的潜在能力结构。分析揭示了塑造规划性能的两个主要维度:操作推理,即评估局部动作适用性和即时状态转换的能力,以及结构枚举,即推理目标可达性和地标结构的能力。操作推理在模型扩展和更长的推理轨迹下得到改善,而结构枚举相对不敏感。我们的发现推动了对大语言模型规划的能力层面评估,将重点从模型是否整体改进转移到哪些规划能力在何种条件下以及为何得到改进。

英文摘要

When LLMs exhibit uneven performance across planning tasks, these gaps are often attributed to task difficulty. We argue that this explanation is incomplete, as task-level variation may reflect distinct latent planning competencies rather than differences along a single ability spectrum. We study this question on ACPBench-Hard by evaluating multiple LLM families under varying test-time reasoning budgets and applying a multidimensional item response theory model to uncover the latent competency structure underlying LLM planning. The analysis reveals two principal dimensions that shape planning performance: operational reasoning, the ability to evaluate local action applicability and immediate state transitions, and structural enumeration, the ability to reason about goal reachability and landmark structure. Operational reasoning improving under model scaling and longer reasoning traces, while structural enumeration remains comparatively insensitive. Our findings motivate competency-level evaluation of LLM planning, shifting the focus from whether models improve overall to which planning competencies improve, under what conditions, and why.

Comments19 pages. Keywords: Reasoning, Automated Planning, Item Responses Theory, LLMs as Planner Research Area: NLP and Symbolic Reasoning Research Area Keywords: neurosymbolic, planning in agents, symbolic reasoning Contribution Types: Model analysis & interpretability

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑