arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

规划与修补:用于智能体规划的扩散语言模型

Plan-and-Patch: Diffusion Language Models for Agentic Planning

Syamantak Kumar, Jiang Guo, Hassan Hamad, Hideo Kobayashi, Yi Xiang, Yezhou Yang, Yanjun Qi, Daniele Bonadiman, Jiarong Jiang

arXiv 2610.10786首次发表:更新:

发表机构

University of Texas at Austin; Amazon(得克萨斯大学奥斯汀分校; 亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Plan-and-Patch框架,采用扩散语言模型生成并修补智能体规划,在无特定任务训练的Natural Plan上,其规划修补成功率是自回归规划器的近两倍,且规划生成延迟降低39%-46%。

AI 中文摘要

规划对于长程智能体愈发重要,其成功执行需在多步骤中协调子目标、工具使用及中间结果。然而规划过程中的假设可能被环境推翻,工具可能返回意外结果,或动作执行失败。因此,高效智能体不仅要生成规划,还需对其进行修订。此类修订通常仅影响规划的部分内容,保留前后结构完整。与其重新生成整个规划并带来不必要的改动,修补可基于保留的前缀和后缀,仅重新生成受影响区域。我们提出Plan-and-Patch,这是一种规划-行动框架,其中扩散语言模型(dLLM)通过并行去掩蔽生成结构化、类程序的规划,并通过填充选定区域、固定周围步骤来修补规划。我们对比DreamReasoner-8B和Qwen3-8B作为扩散规划器与自回归(AR)规划器,在无特定任务训练的Natural Plan上,扩散规划器的规划修补成功率(53.7%)几乎是AR规划器(27.0%)的两倍;在智能体基准ALFWorld和TextCraft上进行特定任务训练后,两类规划器的规划生成观测成功率相近,而扩散规划器的平均规划生成延迟较AR降低39%-46%。结果表明,Plan-and-Patch为长程智能体提供了更快规划生成与有效规划修补的框架。

英文摘要

Planning is increasingly important for long-horizon agents, where successful execution requires coordinating subgoals, tool use, and intermediate outcomes over many steps. Yet assumptions made during planning may be invalidated by the environment, tools may return unexpected results, or actions may fail. Effective agents must therefore not only generate plans, but also revise them. Such revisions often affect only part of a plan, leaving the preceding and subsequent structure intact. Rather than regenerate the entire plan and risk unnecessary changes, repair can regenerate the affected region conditioned on the preserved prefix and suffix. We introduce Plan-and-Patch, a plan-and-act framework in which a diffusion language model (dLLM) generates a structured, program-like plan through parallel unmasking and repairs it by filling in selected regions while keeping the surrounding steps fixed. We compare DreamReasoner-8B and Qwen3-8B as diffusion and autoregressive (AR) planners. On Natural Plan without task-specific training, diffusion (53.7%) achieves nearly twice the plan repair success rate of AR (27.0%). After task-specific training on agentic benchmarks, ALFWorld and TextCraft, the planners achieve similar observed success in plan generation, while diffusion reduces mean plan-generation latency by 39-46% relative to AR. Our results show that Plan-and-Patch provides a framework for faster plan generation and effective plan repair in long-horizon agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑