发表机构
ADIA Lab; University of Granada; Luxembourg Institute of Science and Technology; Cornell University; Lawrence Berkeley National Laboratory(ADIA实验室; 格拉纳达大学; 卢森堡科学技术研究院; 康奈尔大学; 劳伦斯伯克利国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出规划即路由方法,通过声明四种规划模式并分派给模式特定执行器,显著缩小LLM智能体的计划声明-执行差距,提升任务成功率,但模式选择仍待解决。
AI 中文摘要
大型语言模型(LLM)使智能体能够通过生成计划然后在环境中执行来解决长时程任务。然而,成功的规划需要两种不同的能力:为任务选择合适的计划并忠实地执行它。现有的规划器-执行器系统可能在任一阶段失败,而最终任务成功本身无法区分选择失败和执行失败。因此,我们研究了计划声明-执行差距,并引入了规划即路由(Planning-as-Routing),其中LLM声明四种规划模式之一:预定义(Predefined)、顺序(Sequential)、层次(Hierarchical)或搜索(Search),一个确定性路由器将任务分派给相应的模式特定执行器。在四个基准测试和三个LLM中,我们发现了三个一致的模式。首先,通用的Plan+ReAct往往无法保留声明的规划结构,尤其是对于较长的计划:在三个基准测试中,只有22%至45%的轨迹保留了该结构,而模式特定的执行器则强制执行预期的结构。其次,规划模式的有效性因环境和模型而异:搜索(Search)在ALFWorld上表现最佳,层次(Hierarchical)在SWE-bench上表现最佳,最强模式在同一基准内可能因模型而异。第三,最大的收益来自执行:模式特定的执行器将ALFWorld上的任务成功率从0.48提高到0.92,在SWE-bench Verified上将任务成功率从0.36提高到0.44,相对于Plan+ReAct。然而,当前的LLM并不能可靠地为每个任务选择最强模式,尽管在一些基准-模型组合中,少量示例可以提高选择能力。总体而言,可靠的智能体规划需要有效的模式选择和忠实的执行:路由显著缩小了执行差距,而任务特定的模式选择仍然是一个开放问题。
英文摘要
Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve declared planning structure, especially for longer plans: across three benchmarks, only (22)--(45%) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from execution: pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench Verified over Plan+ReAct. Current LLMs, however, do not reliably select the strongest mode for each task, although few-shot examples improve selection in some benchmark--model combinations. Overall, reliable agent planning requires both effective mode selection and faithful execution: routing substantially closes the execution gap, while task-specific mode selection remains open.
Comments51 pages, 8 figures