arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OpenEvoShield:面向开放世界多智能体系统攻击的双重非平稳持续防御

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks

Litian Zhang, Chaozhuo Li, Yuting Zhang, Zejian Chen, Bingyu Yan, Qiwei Ye

arXiv 2607.19351首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; Beihang University; Beijing Academy of Artificial Intelligence(北京邮电大学; 北京航空航天大学; 北京人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM-MAS中对手通过通信注入恶意指令的双重动态攻击,提出OpenEvoShield持续防御框架,含不对称速率控制器等组件,实验表明其在多基准和拓扑上优于基线,能检测多数未见攻击且误报率低。

AI 中文摘要

基于大语言模型的多智能体系统(LLM-MAS)越来越多地部署在安全关键应用中,对手通过智能体间通信注入恶意指令来传播有害行为。此类攻击具有双重动态性,与静态威胁不同。现有防御将部署视为封闭世界问题,一旦分布超出训练范围就会迅速退化。我们提出了OpenEvoShield,一种用于LLM-MAS的协同进化持续防御框架。包括不对称速率控制器(M1)、正常边界更新器(M2)、EWC正则化策略集成(M3)和基于能量的多粒度检测器(M4)。实验表明,OpenEvoShield优于静态和持续基线,能检测大多数未见攻击并保持低误报率。

英文摘要

LLM-based multi-agent systems (LLM-MAS) are increasingly deployed in safety-critical applications, where adversaries inject malicious instructions through inter-agent communication to propagate harmful behaviors. Unlike static threats, these attacks are doubly dynamic: adversaries refine injection strategies against deployed defenses while normal-agent behavior drifts with system expansion. Existing defenses treat deployment as a closed-world problem and degrade rapidly once either distribution shifts beyond training coverage. We propose OpenEvoShield, a co-evolutionary continual defense framework for LLM-MAS. An asymmetric rate controller (M1) decouples fast attack-side and slow normal-side learning rates from dual drift signals. A normal-boundary updater (M2) maintains a dynamic behavioral boundary at the slow rate, while an EWC-regularized policy ensemble (M3) fast-adapts without catastrophic forgetting. An energy-based multi-granularity detector (M4) fuses node-, subgraph-, and graph-level evidence to classify novel attacks as out-of-distribution. Experiments over 100 deployment rounds across five benchmarks and four MAS topologies show that OpenEvoShield outperforms static and continual baselines, detecting most previously unseen attacks while keeping false positive rates low.

Comments29 pages, 5 figures, 14 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑