arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01652cs.ROcs.AI

SyncPlan:通过显式同步与自适应修正实现长视 Large Language Model(大语言模型,LLM)协调

SyncPlan: Long-Horizon LLM Coordination with Explicit Synchronization and Adaptive Correction

Shen You, Xiaoming Zhu, Weining Weng, Hefei Mei, Weixuan Wang, Zhongshen Li, Zeji LI, Ye-Wen Wang, Zijun Liao, Juchao Zhuo, Yang Wei, Fuhao Qiu, Siqin Li, Zhenj… 展开作者

Shen You, Xiaoming Zhu, Weining Weng, Hefei Mei, Weixuan Wang, Zhongshen Li, Zeji LI, Ye-Wen Wang, Zijun Liao, Juchao Zhuo, Yang Wei, Fuhao Qiu, Siqin Li, Zhenjie Lian, Danei Gong, Junkai Ji, Xiangtao Li, Qiuzhen Lin, Liang Wang, Ka-Chun Wong

首次发表
浏览论文内容

中文总结 AI 辅助

SyncPlan 是用于长视 LLM 多智能体协调的“规划-执行-修正”框架,通过显式同步、自适应修正及轻量级计划过时检测器优化,在低耗时下实现了超越现有方法的任务成功率。

中文摘要 AI 辅助

基于 LLM 的多智能体协调在动态环境中面临效率与适应性之间的根本权衡。现有方法通常依赖重复的 LLM 调用或多轮通信,以在执行过程中调整决策,这会引入大量延迟,并使协调易受异步进度和环境变化的影响。相反,一次性规划降低了协调开销,但生成的开环计划在行动依赖于其他智能体和环境时,会很快变得过时或失效。我们提出 SyncPlan,这是一种通过显式同步与自适应修正实现长视协调的“规划-执行-修正”框架。给定状态和团队级任务,一个集中式 LLM 协调器在单次规划调用中生成每个智能体的行动链。在执行期间,显式等待原语和死锁检测机制用于强制智能体间及智能体与环境间的依赖关系,而轻量级计划过时检测器会持续评估剩余计划,并在环境变化使其假设失效时触发重新规划。我们还通过 SFT(监督微调)和面向规划的强化学习(RL)优化协调器,使用密集的任务进度和结果级执行反馈。在公开的 Overcooked 基准和复杂的 Honor of Kings 环境上的实验表明,SyncPlan 实现了最先进的任务成功率,同时与现有基于 LLM 的协调器相比,其挂钟运行时间不到 0.05%。代码和数据集将公开提供。

英文摘要

LLM-based multi-agent coordination faces a fundamental trade-off between efficiency and adaptivity in dynamic environments. Existing approaches typically rely on repeated LLM invocations or multi-round communication to adapt decisions during execution, introducing substantial latency and making coordination vulnerable to asynchronous progress and environmental changes. Conversely, one-shot planning reduces coordination overhead but produces open-loop plans that can quickly become stale or fail when actions depend on other agents and the environment. We introduce SyncPlan, a plan-execute-correct framework for long-horizon coordination through explicit synchronization and adaptive correction. Given the state and team-level task, a centralized LLM coordinator generates per-agent action chains in a single planning call. During execution, explicit wait primitives and deadlock detection enforce inter-agent and agent-environment dependencies, while a lightweight Plan Staleness Detector continuously assesses the remaining plan and triggers replanning when environmental changes invalidate its assumptions. We further optimize the coordinator through SFT and planning-oriented RL with dense task progress and outcome-level execution feedback. Experiments on the public Overcooked benchmark and the complex Honor of Kings environment show that SyncPlan achieves state-of-the-art task success rates while using less than 0.05% of the wall-clock runtime compared with existing LLM-based coordinators. Code and datasets will be made publicly available.

↑