arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分解带来完整性,而非收益

Decomposition Buys Integrity, Not Yield

Rong He

arXiv 2609.17464首次发表:更新:

AI 中文总结

本研究通过理论模型和真实生产数据证明,多智能体任务分解虽能提升完整性,但会降低收益,并量化了深度、对齐成本及委派时机的影响。

AI 中文摘要

多智能体系统将任务分解到一棵智能体树上,并用一些经验法则来证明这种分解的合理性:更小的上下文、更清晰的分离、并行性。我们探究这种分解对叶子节点发现的信息有多少能到达根节点的影响。我们将分解建模为一棵树,其中接收 $b$ 个条目的智能体以概率 $r(b)$ 保留其中任何一个。如果 $r(b)=1/b$,那么对于任何任务规模和任何树形结构,每棵树恰好传递一个发现;我们在20,000棵随机不规则树上验证了这一结果,精度达到 $2.4 \times 10^{-15}$。如果 $r(b)=Cb^{-\delta}$,那么一棵深度为 $k$、包含 $N$ 个发现的树会产生 $C^k N^{1-\delta}$ 个结果:任务规模和架构分离,架构每层仅贡献 $C \le 1$,因此对于收益而言,扁平结构是最优的,并且任何智能体排列都无法逃脱指数 $\delta$。在600条生产级深度研究轨迹中,通过三种不共享失败模式的识别方法,测得 $\delta = 0.34$ [0.30, 0.38]。在条目边界来自工具而非文本启发式、且 $b=1$ 出现550次的跳数中,$C = 0.571$ [0.527, 0.615] 是观测值而非外推值,基于16,082次跳数。层级还会带来对齐成本:在1,012条带注释的多智能体轨迹中,每十六条简报中就有一条偏离目标,得到 $\mu = 0.939$,每层惩罚为 $C\mu = 0.536$。深度在另外两个维度上也有代价。根上下文是唯一持久的状态,也是唯一无法廉价遗忘的状态,深度将其暴露从 $N$ 个条目减少到 $N^{1/k}$。深度也更便宜:生产级扁平智能体的计费为 $N^{1.39}$,而不是追加式上下文所预测的 $N^2$,并且在相同花费下,两层结构在403个发现处超过扁平结构。在我们测量的每个参数上,模型显示0.7%到11.3%的生产会话值得委派,而实际委派的比例为7.8%。对743,819次生产工具调用的风险模型发现,委派并非响应于上下文填满,而是一种开局动作。

英文摘要

Multi-agent systems split a task across a tree of agents and justify the split with folklore: smaller contexts, cleaner separation, parallelism. We ask what the split does to how much of what the leaves discover reaches the root. Model a decomposition as a tree in which an agent handed $b$ items keeps any one with probability $r(b)$. If $r(b)=1/b$, every tree delivers exactly one finding, for every task size and every shape; we verify this to $2.4 \times 10^{-15}$ on 20,000 random irregular trees. If $r(b)=Cb^{-δ}$, a depth-$k$ tree over $N$ findings yields $C^k N^{1-δ}$: task size and architecture separate, and architecture contributes only $C \le 1$ per level, so flat is optimal for yield and no arrangement of agents escapes the exponent $δ$. On 600 production deep-research traces $δ= 0.34$ [0.30, 0.38], by three identifications that do not share a failure mode. At a hop where item boundaries come from the tool rather than a text heuristic, and where $b=1$ occurs 550 times, $C = 0.571$ [0.527, 0.615] is observed rather than extrapolated, over 16,082 hops. A tier also costs alignment: on 1,012 annotated multi-agent traces one brief in sixteen goes off-target, giving $μ= 0.939$ and a per-tier penalty $Cμ= 0.536$. Depth is bought on two other axes. The root context is the only state that persists and the only one that cannot cheaply forget, and depth cuts its exposure from $N$ items to $N^{1/k}$. Depth is also cheaper: production flat agents bill as $N^{1.39}$, not the $N^2$ an append-only context predicts, and at equal spend two tiers overtake flat at 403 findings. Across every parameter we measured the model says 0.7% to 11.3% of production sessions are worth delegating, against 7.8% that do. A hazard model on 743,819 production tool calls finds that delegation does not respond to a filling context and is instead an opening move.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑