发表机构
Shanghai Artificial Intelligence Laboratory; Fudan University; Shanghai Jiao Tong University; The Hong Kong University of Science and Technology(上海人工智能实验室; 复旦大学; 上海交通大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出SHE框架,将LLM智能体的安全管控机制分解为四个构件并引入归因引导演化循环,在Agent-SafetyBench上使攻击成功率降低3.1倍,还具备泛化与跨模型迁移能力。
AI 中文摘要
大语言模型(LLM)智能体的安全性不仅取决于模型权重,还取决于管理上下文、记忆、工具、权限和运行时控制的智能体安全管控机制(harness)。现有安全机制通常将安全管控机制视为固定的部署产物,限制了其应对新出现风险的演化能力。此外,安全管控机制各组件间的耦合功能模糊了安全责任归属,使得局部演化难以开展。我们提出安全管控机制演化(Safety Harness Evolution,SHE)框架,该框架从rollout轨迹中学习演化的安全边界。SHE将安全管控机制分解为具有明确安全责任的四个构件,包括系统提示词(System Prompt)、规则库(Rule Bank)、安全记忆(Safety Memory)和工具策略(Tool Policy),为局部演化定义了清晰的功能边界。基于该分解,SHE引入了归因引导的演化循环,将轨迹失败转化为结构化诊断,学习构件特定的边界优化,并通过安全-效用验证选择演化后的安全管控机制。在Agent-SafetyBench上的实验表明,SHE通过安全管控机制演化有效提升了安全性,与静态SafeHarness相比,实现了3.1倍的攻击成功率(ASR)降低,同时也提升了良性效用。演化后的安全管控机制进一步在保留的AgentHarm基准上对未见过的风险具有泛化能力,且无需额外演化即可跨智能体模型迁移。
英文摘要
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harness as a fixed deployment artifact, limiting their ability to evolve with emerging risks. Moreover, coupled functions across harness components obscure safety responsibility attribution, making localized evolution difficult. We propose Safety Harness Evolution (SHE), a framework that learns evolving safe boundaries from rollout trajectories. SHE decomposes the harness into four artifacts with explicit safety responsibilities, including the System Prompt, Rule Bank, Safety Memory, and Tool Policy, defining clear functional boundaries for localized evolution. Based on this decomposition, SHE introduces an attribution-guided evolution loop that converts trajectory failures into structured diagnoses, learns artifact-specific boundary refinements, and selects evolved harnesses through safety-utility validation. Experiments on Agent-SafetyBench demonstrate that SHE effectively enhances safety through harness evolution, achieving a 3.1x ASR reduction compared with static SafeHarness, while also improving benign utility. The evolved harness further generalizes to unseen risks on the held-out AgentHarm benchmark and transfers across agent models without additional evolution.
CommentsProject: https://github.com/RainbowQTT/SHE