发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出两阶段智能体轨迹审计框架,结合规则与LLM审计,在OpenAgentSafety上减少审计次数和令牌使用,同时保持高检测率。
AI 中文摘要
由大型语言模型(LLM)驱动的AI智能体能够执行复杂任务,但可能有意或无意地对其运行的系统造成损害。现有的智能体监控方法依赖于基于规则的护栏或基于LLM的轨迹审计。然而,基于规则的护栏可能通过混淆技术被绕过,并可能遗漏超出其预定义规则的有害行为,而对每个动作应用LLM进行审计则成本高昂。我们提出了一种两阶段的智能体轨迹审计框架。第一阶段使用单事件和轨迹序列规则来选择待检查的待定动作;第二阶段在执行前,使用LLM审计智能体在智能体先前轨迹的上下文中检查每个选定的动作。我们使用训练数据联合优化门控规则和审计指令,使框架能够适应复杂的智能体行为,而不是仅依赖预定义规则。在公共基准OpenAgentSafety上,我们的框架将每次运行的平均LLM审计次数从8.15次减少到2.33次,令牌使用量从47.8k减少到14.6k,检测率为72.8%,而审计每个动作时的检测率为81.5%。在两个模拟多智能体案例研究中,该框架标记了所有恶意轨迹,同时将审计令牌使用量减少了80%以上。
英文摘要
AI agents powered by large language models (LLMs) can perform complex tasks but may harm the systems they operate in, either intentionally or unintentionally. Existing agent monitoring approaches rely on rule-based guardrails or LLM-based trace auditing. However, rule-based guardrails can be bypassed through obfuscation and may miss harmful actions beyond their predefined rules, whereas applying an LLM to audit every action is costly. We present a two-stage agent trace auditing framework. The first stage uses single-event and trace-sequence rules to select pending actions for inspection; the second uses an LLM audit agent to examine each selected action in the context of the agent's preceding trace before execution. We jointly refine the gate rules and audit instructions using training data, allowing the framework to adapt to complex agent behaviors rather than relying solely on predefined rules. On the public benchmark OpenAgentSafety, our framework reduces the average number of LLM audits from 8.15 to 2.33 per run and token usage from 47.8k to 14.6k, with a detection rate of 72.8\% compared with 81.5\% when every action is audited. In two simulated multi-agent case studies, the framework flags all malicious traces while reducing audit token usage by more than 80\%.