arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32802cs.AIcs.CLcs.LG

可重推导性决定分阶段智能体流水线在上游故障后恢复什么

Re-derivability Decides What a Staged Agent Pipeline Recovers After an Upstream Fault

Tianqi Bu, YuXuan Peng, Junteng Tu, Henghui Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出可重推导性概念,解释上游故障对分阶段智能体流水线的代价,证明重新锚定原始问题可廉价恢复性能,且首个阶段修复最有效。

中文摘要 AI 辅助

一个变量决定了上游故障对分阶段语言模型智能体流水线的代价:可重推导性,即一个阶段所需的信息中有多少可以从原始问题中重建。将检查智能体锚定在该问题上,在四个禁用思考的开源骨干模型上,相比盲检查智能体,其价值从+0.608 [+0.517, +0.700] 提升至+0.358,而在两个Qwen骨干模型上,盲检查智能体完全不改变任何项目。该头对头比较是探索性的。一个确定性故障进入第一阶段,我们将原始问题重新暴露给下游阶段中的$k = 0,\dots,3$个阶段,保持智能体、项目、故障和拓扑结构不变,每个臂在温度为零时使用120个gsm_hard项目。故障下的准确率在四个骨干模型中的全部四个上均有所提升,从+0.233到+0.392,最大的Holm校正$p$值为$2.1\times10^{-6}$。一项注册的剔除测试排除了令牌的影响。将每个单词置空保持词槽固定,保留率在四个骨干模型中的全部四个上随可见比例变化,在主要模型上从0.221攀升至0.692。这两个系列是确认性的,而这里的其他一切均为探索性的。交互作用在注册流水线下的四个骨干模型中的两个上排除零,在三阶段流水线下的四个骨干模型中的全部四个上排除零,在完整消息历史下的四个骨干模型中的三个上排除零,主要模型为+0.317。在Llama-3.1-8B上,故障在任何剂量下都不带来可检测的代价,因此其他三个模型承担了所有关于故障代价的主张。可重推导性还决定了架构的代价,我们测量的任何分解都无法可靠地胜过一次直接调用。在未注入故障的情况下,注册流水线以-0.267、-0.125和-0.317的差距输给该调用,而在Phi-4上读取到+0.058,$p = 0.118$,该测试未能将其与零区分开。有效的修复是廉价且前置的:在Qwen3-14B上,第一个重新锚定的阶段以每个项目+59.8个令牌的代价换取+0.394的匹配保留率,而后续阶段则一无所获。

英文摘要

One variable sets what an upstream fault costs a staged pipeline of language-model agents: re-derivability, how much of what a stage needs it can rebuild from the original problem. Grounding an inspector agent in that problem is worth +0.608 [+0.517, +0.700] to +0.358 over a blind one on four open-weight backbones served with thinking disabled, and on the two Qwen backbones the blind inspector changes no item at all. That head-to-head is exploratory. One deterministic fault enters the first stage, and we re-expose the original problem to $k = 0,\dots,3$ of the downstream stages with agents, items, fault and topology held fixed, on 120 gsm_hard items per arm at temperature zero. Accuracy under fault rises on four of four backbones, from +0.233 to +0.392, the largest Holm-adjusted $p$ being $2.1\times10^{-6}$. A registered kill test rules out tokens. Blanking every word holds the word slots fixed, and retention tracks the visible fraction on four of four, climbing from 0.221 to 0.692 on the primary. Those two families are confirmatory and everything else here is exploratory. The interaction excludes zero on two of four backbones under the registered pipeline, four of four under a three-stage pipeline, and three of four under full message history, the primary at +0.317. On Llama-3.1-8B the fault carries no detectable cost at any dose, so the other three carry every claim about what a fault costs. Re-derivability also sets what the architecture costs, and no decomposition we measured reliably beats one direct call. With no fault injected the registered pipeline loses to that call by -0.267, -0.125 and -0.317, and on Phi-4 reads +0.058 at $p = 0.118$, which the test fails to separate from zero. The repair that works is cheap and front-loaded: the first re-grounded stage buys +0.394 of matched retention for +59.8 tokens per item on Qwen3-14B, and the stages after it buy nothing.

补充信息

↑