发表机构
SAI International School; University of Massachusetts Amherst; SP Jain School of Global Management(SAI国际学校; 马萨诸塞大学阿默斯特分校; SP Jain全球管理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究使用工具的智能体的调用层面可靠性,发现其早期错误会导致后续能力下降,提出条件状态评分的补救措施以修正评分规则的偏差。
AI 中文摘要
使用工具的智能体存在两类失效情况:选择错误的工具,或形成错误的论证,且任意一类的早期失效都可能悄无声息地破坏后续所有内容。我们在无污染的多步骤任务(深度为1-8)上,针对五个开放权重模型,在干净的教师强制上下文和模型自身的自由运行上下文两种场景下,测量了可区分上述两类失效的正确调用率。当深度达到6时,模型自身干净上下文能力的约70%会因自身早期错误而丧失(L6值为0.686、0.684)。我们的核心发现与测量方式本身相关:在针对固定黄金轨迹的精确匹配评分下,传播模型的严重程度和恢复参数不仅难以估计,还会被评分规则固定。严重程度被强制到边界(869个中毒步骤中0个正确);恢复能力在结构上不可观测(580个中毒步骤中0个返回正轨,而随机情况下预期为0.0058)。两者都源于同一机制:偏离正轨后,黄金值由模型从未见过的工具常量生成,属于模型无法推导的信息。即便如此,拟合运行仍会为一个精确等于1.000的量返回0.92和0.73——这是评分规则已确定的参数的自信数值。我们提出了一种机制及补救措施:条件状态评分,可回溯应用于缓存的完成结果且无额外成本,该措施将严重程度解绑定到排除零的内部估计值(+0.149、+0.316)。
英文摘要
Tool-using agents fail two ways: choosing the wrong tool, or forming wrong arguments, and an early failure of either kind can silently corrupt everything downstream. We measure a correct-invocation rate that separates the two, under both a clean teacher-forced context and the model's own free-running context, on five open-weight models over contamination-free multi-step tasks (depths 1-8). By depth 6, roughly 70% of a model's own clean-context capability is lost to its own earlier mistakes (L6 = 0.686, 0.684). Our central finding concerns the measurement itself. Under exact-match scoring against a fixed gold trajectory, a propagation model's severity and recovery parameters are not merely hard to estimate - they are fixed by the scoring rule. Severity is forced to its boundary (0 of 869 poisoned steps correct); recovery is structurally unobservable (0 of 580 poisoned steps returned on-track, against an expected 0.0058 by chance). Both follow from one mechanism: post-divergence, the gold value is generated by tool constants the model never sees, so it is information the model cannot derive. A fit run anyway returns 0.92 and 0.73 for a quantity that is exactly 1.000 - confident numbers for a parameter the scoring rule already determined. We give the mechanism and a remedy, conditional-on-state scoring, applied retrospectively to cached completions at zero additional cost, which un-pins severity to interior estimates excluding zero (+0.149, +0.316).