发表机构
Soongsil University; Yonsei University; Jeju National University(崇实大学; 延世大学; 济州国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文定义了针对自进化智能体自主技能进化流程的EvoSkill注入威胁模型,提出SARGE红队测试框架,构建相关基准并验证其可诱导恶意技能形成与持久激活,凸显能力损坏风险。
AI 中文摘要
基于大语言模型(LLM)的智能体系统日益采用基于技能的架构,以降低重复推理成本并提升任务执行的稳定性与效率。近期研究提出了自进化智能体,这类智能体可从过往经验中自主生成、优化并复用技能,以实现能力的持续进化。然而,自主技能进化引入了新的攻击面,恶意能力会被生成、存储并作为合法技能复用。本文将EvoSkill注入定义为针对自进化智能体自主技能生成与进化流程的威胁模型。我们进一步提出SARGE(自进化智能体中自主技能生成与进化的红队测试框架),该框架通过迭代生成、升级与强化交互来评估此威胁模型。为支撑该框架,我们构建了EvoSkillBench,这是一个用于诱导自进化智能体形成恶意技能的恶意交互轨迹基准数据集,并引入了EvoSkillSafetyBench,这是一个用于评估注入的恶意技能是否会被后续检索并激活为有害行为的攻击后基准。我们的评估表明,SARGE可诱导恶意技能形成,且注入的技能会被持久存储并反复激活,凸显了持久能力损坏的风险。
英文摘要
LLM-based agent systems increasingly adopt skill-based architectures to reduce repetitive reasoning costs and improve stable, efficient task execution. Recent studies propose self-evolving agents that autonomously generate, refine, and reuse skills from past experiences to enable continuous capability evolution. However, autonomous skill evolution introduces a new attack surface in which malicious capabilities are generated, stored, and reused as legitimate skills. In this paper, we define EvoSkill Injection as a threat model targeting the autonomous skill generation and evolution pipeline of self-evolving agents. We further propose SARGE (Red-teaming Autonomous Skill Generation and Evolution in self-evolving agents), a red-teaming framework for evaluating this threat model through iterative generation, escalation, and reinforcement interactions. To support our framework, we construct EvoSkillBench, a benchmark dataset of malicious interaction trajectories for inducing malicious skill formation in self-evolving agents, and introduce EvoSkillSafetyBench, a post-attack benchmark for evaluating whether injected malicious skills are subsequently retrieved and activated as harmful behaviors. Our evaluation shows that SARGE induces malicious skill formation and that injected skills are persistently stored and repeatedly activated, highlighting the risk of persistent capability corruption.
CommentsAccepted to EMNLP 2026