arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06001cs.SEcs.AI

AgentSpy:让AI智能体行为可观察

AgentSpy: Making AI Agent Behavior Observable

Christoph Bühler, Matteo Biagiola, Luca Di Grazia, Guido Salvaneschi

首次发表
浏览论文内容

中文总结 AI 辅助

AgentSpy通过外部监控记录智能体及其子进程的系统调用和网络流量,支持一致性与安全性分析,在77个任务上验证了可靠性,并检测出四类攻击且无误报。

中文摘要 AI 辅助

基于大型语言模型(LLM)构建的AI智能体会运行shell命令、读写文件并访问网络,通常拥有用户的权限。然而,智能体在执行过程中的行为难以理解:测试仅对结果进行断言,而智能体的轨迹记录仅包含智能体自身报告的内容,这可能遗漏其子进程执行的行为。我们提出AgentSpy,一种从智能体外部观察智能体的方法。AgentSpy在隔离环境中运行智能体,该环境由声明式规范配置,并记录智能体及其执行的每个进程的系统调用和网络流量。基于此监控,AgentSpy支持两类分析:一致性分析,用于衡量义务,即智能体执行应做的事情;以及安全性分析,用于检查禁令,即智能体执行绝不能做的事情。我们为每类分析实例化了一个分析。可靠性分析使用规则根据智能体使用的环境资源来总结每次运行:它执行的命令、访问的文件以及联系的主机。安全性分析将确定性规则应用于执行的系统调用。对于可靠性,我们使用codex harness和三个最近的LLM在77个任务上评估了AgentSpy,每个任务执行三次。同一任务的重复运行集合在92.2%的比较中比包含另一任务运行的集合更相似。在所有三次运行均通过基于结果的测试的任务中,智能体在18%的情况下执行与任务无关的活动,在7%的情况下读取评分文件,在17%的情况下未使用开发者的指导。对于安全性,AgentSpy的通用规则检测了我们考虑的五类攻击中的四类,在50次运行中没有误报。

英文摘要

AI agents built on large language models (LLMs) run shell commands, read and write files, and reach the network, typically with their user's privileges. However, what an agent does during an execution is difficult to understand: tests assert on the result, and the agent's trajectory records only what the agent reports about itself, which may omit behavior executed by its subprocesses. We present AgentSpy, an approach that observes an agent from outside the agent. AgentSpy runs the agent in an isolated environment, configured by a declarative specification, and records the system calls and network traffic of the agent and of every process it executes. Based on this monitoring, AgentSpy supports two families of analyses: conformance analyses, which measure obligations, i.e., what an agent execution should do, and safety analyses, which check prohibitions, i.e., what an agent execution must never do. We instantiate one analysis of each family. The reliability analysis uses rules to summarize each run by the environment resources the agent uses: the commands it executed, the files it accessed, and the hosts it contacted. The security analysis applies deterministic rules to the system calls of an execution. For reliability, we evaluated AgentSpy on 77 tasks with the codex harness and three recent LLMs, executing each task three times. Sets of repeated runs of the same task are more similar than sets that include runs of another task in 92.2% of the comparisons. Among tasks for which all three runs pass outcome-based tests, the agent performs task-unrelated activities in 18% of the cases, reads the grading files in 7%, and does not use the developers' guidance in 17%. For security, the generic rules of AgentSpy detect four of five attack categories we considered, with no false positives across 50 runs.

发表机构

  • University of St. Gallen(圣加仑大学)
  • Università della Svizzera italiana (USI)(瑞士意大利语区大学)

机构由 AI 辅助整理,请以论文原文为准。

↑