委托危险带:为何中等能力的子代理过度信任继承的陈旧状态
The Delegation Danger Band: Why Mid-Capability Sub-Agents Over-Trust Inherited Stale State
- PayPal AI(贝宝人工智能)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究揭示中等能力子代理因过度信任继承的陈旧状态而表现下降,形成“危险带”,并提出选择性交接策略以缓解该问题。
AI中文摘要:
智能体框架越来越多地通过分叉子代理来委派工作;一个常见的默认设置是让子代理继承父代理的完整工作上下文。我们测量了继承状态的影响如何随能力变化,其中$C_m$表示干净的分叉新状态下的准确率。我们在同族阶梯(Qwen3 0.6/1.7/4/8B)上,针对一个冻结的、封闭集、动作评分的基准,比较了3种继承策略:重置(分叉新状态:仅基础证据)、选择性(精心策划的交接:+有用的先前结论)和完整(隐式分叉:+有用的结论和$d$份被取代结论的副本)。每个任务都可以从基础证据中解决,因此性能损失可归因于对陈旧状态的依赖。(1) 在2个合成原语以及MuSiQue和HotpotQA上,对被取代状态的遵从度随测量能力$C_m$急剧下降(斜率的置信区间CI在每个家族中均排除零)。(2) 在Qwen3合成阶梯上,净继承危害呈现非单调模式:中等能力模型(Qwen3-1.7B)是净危害的统计显著局部最小值,低于其分叉新状态基线($\Delta(32)=-0.19$ [-0.25, -0.12])及其两个邻居,而最弱的模型保持接近中性,最强的模型保持稳健。我们将这个有害能力范围称为危险带。模型内的计数难度扫描表明,即使在匹配的$C_m$下,效果也取决于模型类别,而实时的父到子分叉重现了中等模型的危害。(3) 精心策划的选择性交接在所有3个数据集上比完整策略提高了平均准确率,在带内模型上提升最大,而固定阈值的路由器在其他数据集上失败;一个可迁移的路由器需要预测重用收益与陈旧上下文惩罚之间的平衡。该基准是冻结的并带有版本哈希。
英文摘要:
Agent frameworks increasingly delegate work by forking sub-agents; a common default makes the child inherit the parent's full working context. We measure how the effect of inherited state changes with capability, where $C_m$ denotes clean fork-fresh accuracy. We compare 3 inheritance policies: Reset (fork fresh: base evidence only), Selective (curated handoff: + the useful prior conclusion), and Full (implicit fork: + the useful conclusion and $d$ copies of a superseded conclusion) over a same-family ladder (Qwen3 0.6/1.7/4/8B) on a frozen, closed-set, action-scored benchmark. Every task is solvable from the base evidence, so performance loss can be attributed to reliance on stale state. (1) Deference to superseded state falls sharply with measured capability $C_m$ (the slope's confidence interval, CI, excludes zero on every family) across 2 synthetic primitives plus MuSiQue and HotpotQA. (2) On the Qwen3 synthetic ladder, net inheritance harm follows a nonmonotone pattern: a mid-capability model (Qwen3-1.7B) is a statistically significant local minimum of net harm, falling below its fork-fresh baseline ($Δ(32)=-0.19$ [-0.25, -0.12]) and both neighbors, while the weakest model stays near-neutral and the strongest models stay robust. We call this harmful capability range a danger band. A within-model counting-difficulty sweep shows that the effect depends on model class even at matched $C_m$, and a live parent-to-child fork reproduces the mid-model harm. (3) Curated Selective handoff improves average accuracy over Full on all 3 datasets, largest at the in-band model, while the fixed-threshold capability router fails on the other datasets; a transferable router would need to predict the balance between reuse benefit and stale-context penalty. The benchmark is frozen and version-hashed.