arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentAudit:一个用于AI智能体全生命周期可信度评估的开放、可扩展框架

AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents

Shrey Nag, Sachita, Abhishek Kumar Singh, Lipi Goel, Rajeshwar Singh Janwar

arXiv 2609.09875首次发表:更新:

发表机构

TIET; GTBIT; IGDTUW; Ministry of Electronics and Information Technology (MeitY), Government of India(塔帕尔工程技术学院; 古鲁·特格·巴哈杜尔理工学院; 英迪拉·甘地德里妇女技术大学; 印度政府电子和信息技术部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AgentAudit是一个开放可扩展的框架,通过评估执行轨迹的十个维度并归因失败,实现AI智能体全生命周期的可信度评估,实验显示不同模型信任分数差异显著。

AI 中文摘要

现有的评估框架大多只评估AI智能体的某一部分,例如任务完成度(AgentBench)或安全鲁棒性(AgentDojo、ASB),而非规划、工具选择、工具执行、记忆和推理的完整流程。失败可能发生在任何阶段,然而现有基准很少能识别其确切来源。AgentAudit在十个能力、接地性、安全性和行为维度上评估整个执行轨迹,即指令完整性、规划器、记忆、工具选择、工具调用、工具正确性、对齐、工具忠实性、安全性和执行完整性,并结合行为分类和失败归因,以精确定位导致观察到的失败的具体阶段。AgentAudit可以评估任何基于LLM的AI智能体,因为它附加到智能体上而非替换它。它只读取记录的执行轨迹,不干扰智能体的运行方式,因此对智能体的内部实现不施加任何约束。我们评估了五个语言模型(OpenAI GPT-5、Claude Sonnet 5、Sarvam 105B、Llama 3.3 70B和Gemini 2.5 Flash),涵盖九个能力和对抗性任务。Claude Sonnet 5和GPT-5获得了最高的平均综合信任分数(分别为95.1和80.6,满分100),而Sarvam 105B、Llama 3.3 70B和Gemini 2.5 Flash则大幅落后(分别为57.6、45.7和22.6)。所有轨迹均由一个固定的评判模型评分,该模型本身也是被评估模型之一,这是第七节E部分讨论的一个局限性。更重要的是,具有相似任务完成行为的模型在可信度上可能差异巨大,因为几个非前沿模型在对抗性任务中反复被分类为“不安全合规”,而不仅仅是失败,这是通过/失败基准无法揭示的区别。

英文摘要

Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion (AgentBench) or security robustness (AgentDojo, ASB), rather than the complete pipeline of planning, tool selection, tool execution, memory and reasoning. Failures can occur at any stage, yet existing benchmarks rarely identify their precise source. AgentAudit evaluates the entire execution trace across ten capability, grounding, security and behavioural dimensions, namely instruction integrity, planner, memory, tool selection, tool invocation, tool correctness, alignment, tool faithfulness, security and execution integrity, combined with behavioural classification and failure attribution to pinpoint the exact stage responsible for an observed failure. AgentAudit can evaluate any LLM-based AI agent, since it attaches to the agent instead of replacing it. It reads only the recorded execution trace and does not interfere with how the agent runs, so it places no constraint on the agent's internal implementation. We evaluate five language models (OpenAI GPT-5, Claude Sonnet 5, Sarvam 105B, Llama 3.3 70B and Gemini 2.5 Flash) across nine capability and adversarial tasks. Claude Sonnet 5 and GPT-5 obtain the highest mean Composite Trust Scores (95.1 and 80.6 out of 100, respectively), while Sarvam 105B, Llama 3.3 70B and Gemini 2.5 Flash trail substantially (57.6, 45.7 and 22.6). All traces were scored by a single fixed judge model, which was itself one of the evaluated models, a limitation discussed in Section VII.E. More importantly, models with similar task-completion behaviour can diverge sharply in trustworthiness, as several non-frontier models are repeatedly classified Unsafe_Compliance on adversarial tasks rather than merely failing them, a distinction that pass/fail benchmarks cannot surface.

Comments23 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑