ToolHazard:用于基于大语言模型智能体的安全评估与对齐的可扩展对抗环境
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
- National Engineering Research Center for Software Engineering, Peking University(北京大学软件工程国家工程研究中心)
- Weixin AI, Tencent Inc.(腾讯微信AI)
- Harbin Institute of Technology(哈尔滨工业大学)
- Beijing University of Posts and Telecommunications(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对集成外部工具的LLM智能体的安全漏洞问题,提出可扩展对抗环境合成框架ToolHazard,构建ToolHazard-Bench测试智能体,生成的对齐数据可提升智能体安全性且保留任务效用。
AI中文摘要:
集成外部工具的大语言模型(LLM)智能体易受嵌入环境状态中的间接提示注入攻击。然而,现有研究大多依赖手动实现或复用的环境、基于随机LLM的工具模拟以及预定义的注入位置,限制了跨更广泛领域的可扩展安全研究。为填补这一空白,我们提出ToolHazard,这是一种可扩展的对抗环境合成框架,可减少人工工程工作量,并支持通过额外的种子领域和计算资源进行扩展。ToolHazard通过环境模拟器、攻击者智能体和用户模拟器,合成可执行的有状态环境,发现可行的注入点并生成特定环境的有效载荷,还能构建基于状态的长周期任务。基于ToolHazard,我们构建了ToolHazard-Bench,用于在复杂工作流和多样化环境攻击下对智能体进行压力测试。实验揭示了智能体存在的大量漏洞,表明注入时机和位置会影响攻击效果。此外,由ToolHazard生成的对齐数据可在保留良性任务效用的同时,提升ToolHazard-Bench和AgentDojo上的安全性。
英文摘要:
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute. Through an Environment Simulator, an Attacker Agent, and a User Simulator, ToolHazard synthesizes executable stateful environments, discovers viable injection points and generates environment-specific payloads, and constructs state-grounded long-horizon tasks. Based on ToolHazard, we build **ToolHazard-Bench** for stress-testing agents under complex workflows and diverse environmental attacks. Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness. Moreover, ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.