SkillDRE:通过预执行和运行时反馈进行智能体技能的双阶段红队演化
SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback
- Beijing University of Posts and Telecommunications(北京邮电大学)
- North China Electric Power University(华北电力大学)
- Chongqing University of Posts and Telecommunications(重庆邮电大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SkillDRE通过双阶段反馈循环(预执行扫描与运行时防御)自动演化恶意技能包,在SkillsBench上对四个模型实现45.28%的平均攻击成功率,超过最强基线40.3%,且最终技能无扫描发现并保留良性任务性能。
AI中文摘要:
智能体技能将指令、可执行代码和任务特定资源打包成可复用的工件,智能体可以利用执行反馈来改进这些工件。同样的机制也使攻击者能够演化恶意技能,使其更有效且更不易被检测。然而,候选技能可能通过预执行扫描,但在运行时防御下未能实现其目标,而修复执行问题的修订可能会引入新的扫描器发现。我们引入了SkillDRE,一个通过双阶段反馈循环演化完整恶意技能包的完全自动化框架。给定一个良性任务及其相关技能,SkillDRE自主构建并验证一个任务条件下的恶意目标和一个可验证的评判规则。然后,它在保持两者固定的同时演化技能实现,并保留合法任务能力。SkillDRE将扫描器引导的演化与运行时引导的细化相结合,后者根据运行时防御下观察到的执行结果提供信息。每次运行时引导的修订都会返回预执行阶段进行重新扫描和进一步优化,然后再重新执行,形成一个跨阶段的闭环。在SkillsBench上对四个受害模型进行评估,SkillDRE实现了45.28%的平均攻击成功率,超过最强基线40.3%,而其最终提交的技能没有收到任何SkillScan发现,并且基本保留了良性任务的性能。这些结果表明,两阶段防御反馈可以作为自适应红队的有用学习信号,而单独评估任一防御阶段可能会错过由此产生的攻击能力。代码可在以下https URL获取。
英文摘要:
Agent skills package instructions, executable code, and task-specific resources into reusable artifacts that agents can improve using execution feedback. The same mechanism also enables attackers to evolve malicious skills, making them more effective and less detectable. However, a candidate skill may pass pre-execution scanning yet fail to realize its target under runtime defenses, while a revision that repairs execution may introduce new scanner findings. We introduce SkillDRE, a fully automated framework for evolving complete malicious skill packages through a dual-stage feedback loop. Given a benign task and its associated skills, SkillDRE autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule. It then holds both fixed while evolving the skill implementation, with preservation of legitimate task capability. SkillDRE combines scanner-guided evolution with runtime-guided refinement informed by execution outcomes observed under runtime defense. Each runtime-guided revision returns to the pre-execution stage for rescanning and further optimization before re-execution, forming a cross-stage closed loop. Evaluated on SkillsBench across four victim models, SkillDRE achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance. These results show that two-stage defense feedback can serve as a useful learning signal for adaptive red teaming and that evaluating either defense stage in isolation can miss the resulting attack capability. Codes is available at https://github.com/whfeLingYu/SkillDRE