arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SquidAgent:明智并行,高效协调

SquidAgent: Parallelize Wisely, Coordinate Efficiently

Yexiong Lin, Shanshan Ye, Yu Yao, Zhen Fang, Bo Han, Tongliang Liu

arXiv 2610.08647首次发表:更新:

发表机构

The University of Sydney; Mohamed bin Zayed University of Artificial Intelligence; University of Technology Sydney; Hong Kong Baptist University(悉尼大学; 穆罕默德·本·扎耶德人工智能大学; 悉尼科技大学; 香港浸会大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多智能体并行执行效率低下的问题,提出SquidAgent,通过令牌预算估算和共享约定块消除重新探索与协调成本,实现2.2倍吞吐量提升。

AI 中文摘要

基于大语言模型的智能体能够解决复杂的多步骤任务,但顺序执行会带来显著的延迟。原则上,将工作并行分配给多个智能体应能带来接近线性的加速。然而,现有的并行多智能体系统往往比单智能体基线运行得更慢。我们将这一差距归因于并行执行带来的两个隐藏成本,而串行智能体则不会产生这些成本。首先,是重新探索成本:并行工作线程在重建编排器已拥有的上下文(如先前决策)时所花费的冗余努力,而这些上下文在串行执行中本会被隐式继承。其次,是协调成本:协调独立生成输出之间不一致性所需的开销。因此,我们推导出一个原则性的决策准则:只有当某层的关键路径成本加上重新探索和协调开销低于相应的串行成本时,该层才应被并行化。虽然这一准则自然以墙钟时间表示,但我们观察到,大语言模型在被要求估计任务持续时间时校准不佳。为解决这一问题,我们转而以预测的输出令牌数来衡量成本,并凭经验发现大语言模型对令牌数的估计比墙钟时间可靠得多。基于这一令牌准则,我们提出了SquidAgent。它在单次规划步骤中估算所有令牌预算,将每个工作线程直接从编排器的会话中分叉以消除重新探索成本,并用预生成的共享约定块取代事后协调,将协调转化为有界的前期成本。然后,一个确定性调度器逐层应用该准则。实验上,SquidAgent相比Claude Code实现了2.2倍的平均吞吐量提升和2.6倍的平均墙钟时间加速,相比最强的多智能体基线实现了2.0倍的吞吐量提升。

英文摘要

LLM-based agents solve complex multi-step tasks, but sequential execution incurs substantial latency. In principle, parallelizing work across multiple agents should yield near-linear speedups. Yet existing parallel multi-agent systems often run slower than a single-agent baseline. We attribute this gap to two hidden costs that parallel execution incurs but a serial agent avoids. First, there is a re-exploration cost: redundant effort spent by parallel workers reconstructing context that the orchestrator already possesses, such as prior decisions, that would otherwise be inherited implicitly in a serial execution. Second, there is an alignment cost: the overhead required to reconcile inconsistencies across independently generated outputs. We thus derive a principled decision criterion: a layer should be parallelized only when its critical-path cost, plus re-exploration and alignment overheads, is lower than the corresponding serial cost. While this criterion is naturally expressed in wall-clock time, we observe that LLMs are poorly calibrated when asked to estimate task duration. To address this, we instead measure cost in predicted output tokens, which we empirically find LLMs can estimate substantially more reliably than wall-clock time. Building on this token-based criterion, we propose SquidAgent. It estimates all token budgets in a single planning step, forks each worker directly from the orchestrator's session to eliminate re-exploration cost, and replaces post-hoc reconciliation with a pre-generated shared convention block that converts alignment into a bounded upfront cost. A deterministic scheduler then applies the criterion layer by layer. Empirically, SquidAgent achieves a 2.2$\times$ mean throughput improvement and a 2.6$\times$ mean wall-time speedup over Claude Code, and a 2.0$\times$ throughput improvement over the strongest multi-agent baseline.

CommentsAccepted at NeurIPS 2026. 37 pages, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑