arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GRASP:基于智能体AI的战略规划生成、修订与评估

GRASP: Generating, Revising, and Assessing for Strategic Planning with Agentic AI

Arunabh Srivastava, Mohammad A., Khojastepour, Srimat Chakradhar, Sennur Ulukus

arXiv 2609.30147首次发表:更新:

发表机构

University of Maryland, College Park; NEC Laboratories America, Inc.(马里兰大学帕克分校; NEC美国实验室公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GRASP是一个策略感知的多阶段规划框架,通过解耦生成、修订和评估模块,显著提升了复杂任务中LLM的规划准确率,并在多任务场景下消除了性能下降。

AI 中文摘要

大型语言模型(LLMs)通常表现出一种性能特征,即随着任务复杂度的增加,其可靠性会下降。我们通过引入GRASP——一个策略感知的多阶段规划框架——来解决为复杂任务生成高质量自然语言可执行计划的挑战。GRASP将规划流程解耦为专门的、上下文隔离的模块:它预编译全局宏观指导方针(GenPlan),在隔离的上下文窗口内探索替代的局部策略(RevPlan),并使用多标准判别器(VerPlan)独立评估轨迹。实证评估表明,GRASP在多个数据集上持续建立了新的最先进水平,相比直接LLM规划器,在Natural Plan日历调度(约12.4%提升)、ZebraLogic(约30.8%提升)和SciBench数学上取得了显著的准确率提升。关键的是,在多任务扩展下——标准规划器会立即遭遇性能崩溃——GRASP完全消除了多任务性能下降的惩罚。在交错的双任务环境中,GRASP相比直接LLM规划器实现了高达16.7%的绝对准确率提升。此外,通过隔离上下文和强制严格的宏观正则化,GRASP以14.5%的差距超越了前沿推理模型(如GPT-5-mini)。

英文摘要

Large Language Models (LLMs) typically exhibit a performance profile where reliability degrades as task complexity increases. We address the challenge of generating high-quality natural language executable plans for complex tasks by introducing $\textbf{GRASP}$, a strategy-aware, multi-stage planning framework. GRASP decouples the planning pipeline across specialized, context-isolated modules: it pre-compiles global macro-guidelines (GenPlan), explores alternative localized strategies within isolated context windows (RevPlan), and independently evaluates trajectories using a multi-criteria discriminator (VerPlan). Empirical evaluations show that GRASP consistently establishes a new state-of-the-art frontier across diverse datasets, yielding substantial accuracy gains over direct LLM planners on Natural Plan Calendar Scheduling ($\sim$12.4$\%$$\uparrow$), ZebraLogic ($\sim$30.8$\%$$\uparrow$), and SciBench Math. Crucially, under multi-task scaling-where standard planners suffer immediate performance collapse-GRASP completely flattens the multi-task degradation penalty. In interleaved dual-task environments, GRASP achieves an absolute accuracy gain of up to 16.7$\%$ over direct LLM planners. Furthermore, by isolating context and enforcing strict macro-regularization, GRASP outperforms frontier reasoning models (such as GPT-5-mini) by a margin of 14.5$\%$.

CommentsAccepted at the Second Workshop for Research on Agent Language Models (REALM) at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑