arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单独无害,联合有害:在自进化大语言模型智能体中利用经验组合

Benign Alone, Harmful Together: Exploiting Experience Composition in Self-Evolving LLM Agents

Bingyu Yan, Xiaoming Zhang, Chaozhuo Li, Ziyi Zhou, Yirui Qi, Litian Zhang

arXiv 2608.01759首次发表:更新:

AI 中文总结

本研究提出EvoBreak攻击,利用自进化LLM智能体的良性经验组合,结合BreakGym生成训练目标,在保持高无害性的同时优于现有攻击,揭示了新的安全风险。

AI 中文摘要

自进化大语言模型智能体通过将交互轨迹提炼为持久经验来提升能力,但该机制引入了新的安全风险:单独无害的经验在跨会话积累和重用时,可能共同削弱智能体的安全边界。现有记忆攻击通常需要直接访问内存或诱导明确的恶意记录,限制了其隐蔽性和适用性。我们提出EvoBreak,一种基于经验条件的序列攻击,通过单独无害的攻击阶段任务和诱导的经验运作。EvoBreak反复观察受害者提炼的经验,识别未覆盖的目标相关需求,并自适应获取补充经验,随后重新构建最终查询以联合激活这些经验。为支持训练,我们引入BreakGym,一种优先结构的合成流水线,生成具有多样依赖结构的可分解安全敏感目标。EvoBreak通过拒绝采样监督微调与提示引导的GRPO进行优化。在自进化框架、受害者主干、预进化领域和安全基准上的实验表明,EvoBreak在保持高无害性的同时,始终优于现有攻击。这些结果揭示,良性经验组合是自进化智能体中持续存在的攻击面。

英文摘要

Self-evolving large language model agents improve their capabilities by distilling interaction trajectories into persistent experiences. Yet this mechanism introduces a new safety risk: experiences that are benign in isolation may jointly weaken an agent's safety boundary when accumulated and reused across sessions. Existing memory attacks typically require direct memory access or induce explicitly malicious records, limiting their stealthiness and applicability. We propose EvoBreak, an experience-conditioned sequential attack that operates through individually benign attack-stage tasks and induced experiences. EvoBreak repeatedly observes the experiences distilled by the victim, identifies uncovered target-relevant requirements, and adaptively acquires complementary experiences before reformulating the final query to activate them jointly. To support training, we introduce BreakGym, a structure-first synthesis pipeline that generates decomposable safety-sensitive targets with diverse dependency structures. EvoBreak is optimized using rejection-sampling supervised fine-tuning and Hint-guided GRPO. Experiments across self-evolving frameworks, victim backbones, pre-evolution domains, and safety benchmarks demonstrate that EvoBreak consistently outperforms existing attacks while maintaining high benignness. These results reveal benign experience composition as a persistent attack surface in self-evolving agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑