arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越手工安全:面向大语言模型智能体的自演化防御

Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents

Jiajun Ruan, Peiyang Li, Yukun Chen, Fengting Li, Chao Feng

arXiv 2608.12977首次发表:更新:

发表机构

University of Minnesota; Ant Group; Tsinghua University; Zhejiang University(明尼苏达大学; 蚂蚁集团; 清华大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对 LLM 智能体的安全威胁,提出自演化运行时防御框架 HARD,将防御开发从人工工程转为自主演化,在保留任务效用的同时提升了安全性能。

AI 中文摘要

大语言模型(LLM)智能体不断扩展的操作能力带来了复杂的安全威胁。运行时防御通过将安全机制集成到智能体执行循环中,已成为缓解这些风险的有效方法。然而,现有的运行时防御严重依赖人工设计的干预措施,缺乏构建和维护的原则性框架。在本研究中,我们首先开发了一种针对运行时防御的 harness 级别的表述,该表述系统地描述了 harness 机制如何支持防御构建,并从 harness 视角提供了对现有运行时防御干预措施的统一视图。基于该表述,我们提出了 HARD(Harness-based Autonomous Runtime Defense Evolution,基于 harness 的自主运行时防御演化),这是一种自演化运行时防御框架,可自动识别合适的干预策略,并基于观察到的失败痕迹迭代改进防御构件。HARD 将运行时防御开发从人工工程转变为自主演化过程,大量实验表明,它在保留良性任务效用的同时,比现有的手工防御提升了安全性能。我们的研究结果表明,自主防御演化是保护已部署 LLM 智能体的有前景的新范式,使智能体能够识别防御弱点并持续改进其保护机制。

英文摘要

The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have emerged as an effective approach to mitigating these risks by integrating security mechanisms into the agent execution loop. However, existing runtime defenses rely heavily on manually designed interventions and lack a principled framework for their construction and maintenance. In this work, we first develop a harness-level formulation of runtime defense that systematically characterizes how harness mechanisms enable defense construction and provides a unified view of existing runtime defense interventions from a harness perspective. Building on this formulation, we propose HARD (Harness-based Autonomous Runtime Defense Evolution), a self-evolving runtime defense framework that automatically identifies appropriate intervention strategies and iteratively improves defense artifacts based on observed failure traces. HARD transforms runtime defense development from manual engineering into an autonomous evolution process, and extensive experiments demonstrate that it improves security performance over existing handcrafted defenses while preserving benign task utility. Our findings highlight autonomous defense evolution as a promising new paradigm for securing deployed LLM agents, enabling agents to identify defense weaknesses and continuously improve their protection mechanisms.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑