SSCBench:评估使用工具的大语言模型智能体的故障注入测试的证据有效性
SSCBench: Evaluating the Evidential Validity of Fault-Injection Tests for Tool-Using LLM Agents
- Jiangnan University(江南大学)
- Nanjing University(南京大学)
- Zhejiang University(浙江大学)
- Wuxi University of Technology(无锡职业技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对使用工具的LLM智能体故障注入评估的证据有效性问题,构建SSCBench协议实例,经实验发现故障采纳的证据条件存在差异,为故障注入评估提供了关键参考。
AI中文摘要:
故障注入正越来越多地用于评估使用工具的大语言模型(LLM)智能体的可靠性。然而,当智能体自身决定执行过程中哪些权威观测值可见时,关于应如何解释故障采纳结果的研究十分有限。本文针对智能体故障注入评估中的这一证据有效性问题展开系统研究。我们开发了一种测量协议,该协议明确规定哪些观测值可以反驳注入的断言,确定这些观测值是否能在受影响事实首次被使用前变得可见,并记录被评估的执行是否实际实现了该条件。我们构建SSCBench作为该协议的实例,并在两个τ-bench环境中对1191次故障执行评估了4种故障算子和5种智能体配置。实验表明,同一已采纳故障案例和智能体配置在不同执行中可实现显著不同的证据条件,且即使支持及时反证主张的群体稀少或不存在,聚合采纳仍可保持明确定义。例如,在44次最终出现反证的采纳运行中,仅有17次在首次使用前收到反证,27次在首次使用后收到。我们还发现,首次错误时机与后续立场修正未必一致,自动轨迹分析可在无法可靠恢复时间诊断所需的首次故障依赖事件的情况下,仍能恢复采纳情况。我们认为,执行实现的证据条件和支持特定主张解释的群体本身就是故障注入评估的一部分,在将采纳解释为使用前反证下的失败之前,应予以报告。
英文摘要:
Fault injection is increasingly used to evaluate the reliability of tool-using LLM agents. However, there has been limited study of how fault-adoption results should be interpreted when the agent itself determines which authoritative observations become visible during execution. In this paper, we present a systematic study of this evidential validity problem in agent fault-injection evaluation. We develop a measurement protocol that specifies what observations can refute an injected assertion, determines whether they can become visible before the affected fact is first used, and records whether the evaluated execution actually realizes this condition. We construct SSCBench as an instantiation of the protocol and evaluate four fault operators and five agent configurations over 1,191 faulted executions in two $τ$-bench environments. Our experiments show that an admitted fault case and agent configuration can realize substantially different evidential conditions across executions, and that aggregate adoption can remain well defined even when the population supporting a timely-counterevidence claim is sparse or absent. For example, among 44 adopted runs in which counterevidence eventually became visible, only 17 received it before first use, while 27 received it afterward. We also find that first-error timing and later stance revision need not coincide, and that automated trajectory analysis can recover adoption without reliably recovering the first faulty-reliance event needed for temporal diagnosis. We argue that the evidential condition realized by an execution and the population supporting a claim-specific interpretation are part of fault-injection evaluation itself and should be reported before adoption is interpreted as failure under pre-use counterevidence.