arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越端到端成功:诊断长视野安全大语言模型智能体的失败

Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents

Wei Shao, Chongzhou Fang, Zuxiong Tan, Zequan Liang, Setareh Rafatirad, Avesta Sasan, Houman Homayoun

arXiv 2608.20563首次发表:更新:

发表机构

University of California, Davis; Rochester Institute of Technology(加利福尼亚大学戴维斯分校; 罗切斯特理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对长视野安全LLM智能体,提出带检查点的诊断方法,发现Gemini 2.5 Flash失败多因未观测到需复用状态,92种子研究显示针对性指导可提升其状态观测率至95.4%,且失败主因随模型代际变化,需诊断失败根源而非仅看总体成功。

AI 中文摘要

长视野安全大语言模型(LLM)智能体必须在多个相互依赖的交互中传递信息和决策,后续行动往往依赖于早期发现的服务、状态或访问权限,这使得最终任务成功的解读变得困难:智能体可能在触及可运用目标能力的节点之前就已失败。本文提出一种诊断方法,该方法通过检查点对安全任务进行检测,将失败划分为能力暴露前和暴露后两类,并利用受控干预测试疑似的上游瓶颈。我们在四类任务中评估该方法,涵盖所发现信息的延迟复用、观测状态的复用、失败策略的恢复以及不确定结果后的决策。针对观测状态复用,检查点分析显示,Gemini 2.5 Flash 的诸多失败发生在模型观测到其后续需复用的状态之前。在一项预先设定的92个随机种子研究中,针对性的协议歧义消除指导使状态观测率从匹配的无指导控制消息下的65.5%提升至95.4%;采用相同设计在Gemini 3.7 Flash上进行测试则产生相反效果,且状态观测不再可靠地预测任务完成情况。这些结果表明,失败的主要来源可能会随模型代际变化,这推动了需诊断长视野安全智能体失败的位置与原因的评估,而非仅依赖总体任务成功。

英文摘要

Long-horizon security LLM agents must carry information and decisions across many dependent interactions, where later actions often depend on services, state, or access discovered much earlier. This makes final task success difficult to interpret: an agent may fail before it ever reaches the point where the capability of interest can be exercised. We present a diagnostic methodology that instruments security tasks with checkpoints, separates failures before and after capability exposure, and uses controlled interventions to test suspected upstream bottlenecks. We evaluate the methodology across four task families involving delayed reuse of discovered information, reuse of observed state, recovery from failed strategies, and decision making after uncertain outcomes. On observed state reuse, checkpoint analysis shows that many Gemini 2.5 Flash failures occur before the model observes the state it is later expected to reuse. In a pre-specified 92-seed study, targeted protocol-disambiguation guidance increases state observation from 65.5\% under a matched non-guidance control message to 95.4\%. Repeating the same design with Gemini 3.7 Flash produces the opposite effect, while state observation no longer reliably predicts task completion. These results show that the dominant source of failure can shift across model generations, motivating evaluation that diagnoses where and why long-horizon security agents fail rather than relying only on aggregate task success.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑