arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

注意力税、交接税:多智能体LLM系统何时有帮助的简约模型

Attention Tax, Handoff Tax: A Stylised Model of When Multi-Agent LLM Systems Help

Akshit Anchan, Nayonika Sen

arXiv 2610.06069首次发表:更新:

发表机构

University of Amsterdam; Northeastern University(阿姆斯特丹大学; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一个简约可靠性模型,通过注意力税与交接税的权衡解释多智能体LLM系统何时优于单智能体,并在账本核对任务上验证了交叉深度预测。

AI 中文摘要

近期关于多智能体LLM系统的研究得出了截然不同的结论:一些结果表明,拥有相同信息和计算资源的单一智能体应优于委派系统,另一些则表明多智能体的收益随任务深度而增长。我们认为,大部分分歧源于对不同瓶颈的建模,并引入了一个围绕两种权衡构建的简约可靠性模型。分解减轻了长上下文的负担,但在智能体之间压缩或传递信息时会产生交接税。冗余从多次采样中获益,但其收益取决于失败在多大程度上是共享的。在加入推理预算、验证和任务结构后,该模型产生了两个交叉条件:一旦通过重置上下文避免的注意力成本超过交接成本,分解就变得更为可取;而当共享失败下限低于单个智能体思考更长时间的误差下限时,在同等预算下并行采样最终更为可取。我们将这些机制与近期的理论和实证结果联系起来。在一个账本核对任务上,我们仅通过单智能体和交接运行测量了上下文退化曲线和交接税。由此,模型将交叉点置于深度10,并预测分解在深度20、50和100时获胜。在步骤级和最终余额准确性上确实如此,且分解系统的成功率(预测从未见过)在每个深度均落在预测率的9个百分点以内。

英文摘要

Recent work on multi-agent LLM systems reaches sharply different conclusions: some results show that a single agent with the same information and compute should dominate a delegated system, others that multi-agent gains grow with task depth. We argue that much of the disagreement comes from modelling different bottlenecks, and introduce a stylised reliability model built around two trade-offs. Decomposition reduces the burden of long contexts but incurs a handoff tax when information is compressed or transferred between agents. Redundancy gains from multiple samples, but its benefit depends on how much their failures are shared. With reasoning budget, verification, and task structure added, the model yields two crossover conditions: decomposition becomes preferable once the attention cost avoided by resetting context exceeds the handoff cost, and parallel sampling at equal budget is eventually preferable when its shared-failure floor lies below the error floor of one agent thinking longer. We connect these regimes to recent theoretical and empirical results. On a ledger-reconciliation task we measure the context-degradation curve and the handoff tax from single-agent and handoff runs alone. From these the model places the crossover at depth 10 and predicts decomposition to win at depths 20, 50, and 100. It does, on step-level and final-balance accuracy, and the decomposed system's success, which the prediction never sees, lands within 9 percentage points of the predicted rate at every depth.

Comments23 pages, 6 figures. Code and data: https://github.com/akshitanchan/attention-handoff-tax

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑