自进化防御:面向LLM智能体的持续安全策略学习
Self-Evolving Defense: Continual Security Policy Learning for LLM Agents
浏览论文内容
中文总结 AI 辅助
提出免训练的自进化防御框架,通过提炼有害轨迹为可复用策略并检索应用,使LLM智能体持续适应新攻击,在多个基准上显著降低提示注入和越狱成功率,同时保持任务效用。
中文摘要 AI 辅助
大型语言模型(LLMs)日益驱动着能够访问敏感信息、使用外部工具并修改软件仓库的智能体。尽管这些能力带来了显著益处,但也引入了诸如越狱、提示注入和易受攻击的代码生成等安全风险。现有防御措施往往需要重新训练,无法适应不断演变的攻击,或仅针对单一威胁模式。为解决这些局限,我们提出了自进化防御(SED),一种免训练框架,它通过将有害智能体轨迹提炼为可复用的安全策略,而无需更新模型权重。通过为未来任务检索相关策略,SED能够持续适应新攻击,同时保留跨攻击场景的知识。为评估SED的有效性,我们在八个涵盖越狱、提示注入和不安全代码生成的基准上,使用三个开源模型(DeepSeek V4 Flash、GLM 5.2和Kimi K3)进行了测试。SED将AGENTDOJO上的定向提示注入成功率降至0.42%,而最佳基线防御为3.7%;在HARMBENCH上将自适应X-TEAMING攻击成功率保持在7.8%,比最佳基线(35.2%)低四倍以上,同时保持了良性任务效用。
英文摘要
Large language models (LLMs) increasingly power agents that access sensitive information, use external tools, and modify software repositories. Although these capabilities offer substantial benefits, they also create security risks such as jailbreaks, prompt injection, and vulnerable code generation. Existing defenses often require retraining, fail to adapt to evolving attacks, or address only a single threat pattern. To address these limitations, we propose Self-Evolving Defense (SED), a training-free framework that distills harmful agent trajectories into reusable security policies without updating model weights. By retrieving relevant policies for future tasks, SED continually adapts to new attacks while retaining knowledge across attack scenarios. To evaluate the effectiveness of SED, we test it with three open-source models (DeepSeek V4 Flash, GLM 5.2, and Kimi K3) on eight benchmarks that span jailbreaks, prompt injection, and insecure code generation. SED lowers targeted prompt-injection success on AGENTDOJO to 0.42%, compared with 3.7% for the best baseline defense, and holds adaptive X-TEAMING attack success on HARMBENCH to 7.8%, more than four times lower than the best baseline at 35.2%, while preserving benign task utility.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。