响亮的失败,安静的失败:使用工具的语言模型智能体中的故障检测与恢复
Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents
- College of Media and Communication Texas Tech University(德州理工大学媒体与传播学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过故障注入测试工具使用语言模型智能体,发现它们过度信任静默失败,仅对显式错误响应,推理模型未改善检测,恢复率受限于随机性。
AI中文摘要:
工具使用智能体通常根据其在工具正常工作时完成任务的情况进行评分。然而实际部署环境并不那么宽容:服务超时、端点消失、参数名称更改,以及结果格式良好但内容错误。先前的研究表明,语言模型过度信任静默失败的工具输出;我们探讨这种过度信任在多轮智能体的故障处理各阶段中如何表现。我们将一个成熟的函数调用基准的可执行环境包裹在故障注入层中,在轨迹中的受控点注入四种类型化故障之一,并记录智能体是否注意到故障、改变计划、恢复任务或重复自身。来自三个模型家族的六种模型(其中一半为推理变体)在24个多步骤任务上运行了1,920次试验。当工具返回显式错误时,智能体在91.3%的试验中将故障视为问题;但当工具返回看似合理的错误值时,该比例仅为58.8%,而在无故障情况下报告问题的比例为26.8%。推理模型并未表现更好:与指令调整的同类模型相比,它们注意到故障更少(-9.3个百分点,p<.001),改变计划更多(+10.4个百分点,p<.001),而恢复率无显著变化(p=.512)。由于智能体具有随机性,同一任务的无故障运行仅有63.3%的时间以相同状态结束;以此基线为对照,只有工具缺失明显降低了恢复率(39.9%),而超时、模式漂移和损坏仍在运行间变异的范围内。故障发生后,智能体连续三次或更多次返回同一工具的情况在高达22.2%的试验中出现,尽管严格相同的重复很少见。一条提示要求智能体检查每个结果并未改善检测。智能体响应错误通道而非工具返回的内容,因此符合预期格式的故障会被忽略。
英文摘要:
Tool-using agents are usually scored on whether they finish a task while the tools work. Deployments are less forgiving: services time out, endpoints disappear, parameter names change, and results come back well formed but wrong. Prior work has shown that language models over-trust tool outputs that fail silently; we ask how that over-trust plays out across the stages of failure handling in multi-turn agents. Wrapping the executable environments of an established function-calling benchmark in a fault-injection layer, we inject one of four typed faults at a controlled point in the trajectory and record whether the agent notices, changes plan, recovers the task, or repeats itself. Six models from three families, half of them reasoning variants, ran 1,920 trials over 24 multi-step tasks. Agents treat a failure as a problem in 91.3% of trials when the tool returns an explicit error, but in 58.8% of trials when it returns a plausible wrong value, against a 26.8% rate of reporting problems when nothing was wrong. Reasoning models are not better placed: paired against instruct siblings, they notice less (-9.3 points, p < .001) and change plan more (+10.4 points, p < .001), and recovery is unchanged (p = .512). Because agents are stochastic, two fault-free runs of the same task end in the same state only 63.3% of the time; against that baseline, only a missing tool clearly lowers recovery (39.9%), while timeouts, schema drift, and corruption stay within run-to-run variation. After a fault, agents return to the same tool three or more times in a row in up to 22.2% of trials, though strictly identical repeats are rare. A prompt line asking the agent to check each result did not move detection. Agents respond to the error channel rather than to the content of what a tool returns, so failures that stay inside the expected format pass through.