发表机构
Vectara, Inc.; Anthropic; OpenAI; Google DeepMind; DeepSeek-AI(Vectara公司; Anthropic公司; OpenAI公司; 谷歌DeepMind; DeepSeek人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型智能体安全可靠行为问题,提出GuardianAgentBench基准测试,含多阶段验证和攻击模式。实验发现模型有两种失败模式,性能随工具集和轮次深度下降。护栏实现优于系统提示防御,证明执行时结构干预可提升安全性且不影响正确行为。
AI 中文摘要
随着大语言模型智能体越来越多地通过访问工具和外部环境进行自主操作,确保其安全可靠的行为变得至关重要。我们提出了GuardianAgentBench(GABench),这是一个在六个领域包含580个场景的基准测试,在LangChain、LlamaIndex和Vectara三个生产就绪框架上进行评估。该基准测试纳入了严格的多阶段验证和五种对抗攻击模式。对六个最先进模型的实验表明,即使是最强配置的整体准确率也仅为74.8%,并揭示了两种不同的失败模式:更强的模型调用所需工具不足,而较弱的模型错误选择和过度调用工具。性能随着工具集大小和顺序轮次深度单调下降,长期规划是更严重的瓶颈。我们的护栏实现始终优于基于系统提示的防御,在误报率仅为0.5%的情况下恢复了19.9%的失败情况。这些结果表明,执行时的结构干预可提高安全性而不干扰智能体的正确行为。
英文摘要
As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavior becomes critical. We present GuardianAgentBench (GABench), a benchmark of 580 scenarios across six domains evaluated on three production-ready frameworks: LangChain, LlamaIndex, and Vectara. The benchmark incorporates rigorous multi-stage validation and five adversarial attack modes. Experiments with six state-of-the-art models reveal that even the strongest configuration achieves only 74.8% overall accuracy and expose two distinct failure regimes: stronger models under-call required tools, while weaker models mis-select and over-call tools. Performance degrades monotonically with both tool-set size and sequential turn depth, with long-horizon planning proving the steeper bottleneck. Our guardrail implementation consistently outperforms system-prompt-based defenses across all models, recovering 19.9% of failures at a false positive rate of just 0.5%. These results demonstrate that execution-time structural intervention improves safety without disrupting correct agent behavior.