arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

防御即技能:为技能增强型智能体演化运行时防护技能

Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents

Xiaofang Yang, Ziqi Miao, Dianbo Sui, Jing Shao, Lijun Li

arXiv 2609.01487首次发表:更新:

发表机构

Shanghai Artificial Intelligence Laboratory; Fudan University; Harbin Institute of Technology, Weihai(上海人工智能实验室; 复旦大学; 哈尔滨工业大学(威海))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出 Defense-as-Skill 防御范式,将运行时防护实现为可安装技能,开发 SkillSonar 防护工具,构建 SCOPE-R 数据集,经演化后可大幅降低攻击成功率,同时保持安全-效用权衡。

AI 中文摘要

技能增强型智能体将可复用技能作为持久运行时上下文加载,提升了任务性能,但也为恶意技能提供了持久渠道以引导未来动作。此类技能可能会泄露秘密、破坏代码、绕过审批,或仅在具体用户任务和工作空间状态使不安全动作显得有用时才为数据 exfiltration 做准备。这使得预安装审查不够充分,需要基于任务的运行时保护。我们提出 Defense-as-Skill,一种将运行时防护本身实现为可安装、可检查和可编辑技能的防御范式。我们的防护工具 SkillSonar 与不可信任务技能并行运行,根据用户的任务边界检查敏感动作,将每个动作路由至允许、重新规划或确认决策,而无需修改底层智能体运行时。为研究该设置,我们构建了 SCOPE-R,一个涵盖 6 个风险家族和 21 个子类的基于任务的数据集,包含 206 个经攻击确认的恶意实例和 43 个良性任务。随后,我们使用运行时防护技能演化在 SCOPE-R 训练子集上改进 SkillSonar,该方法是一种蒙特卡洛树搜索过程,通过 rollout 反馈演化磁盘上的防护技能。在 Claude Code 和 OpenClaw 上,演化后的防护工具大幅降低了攻击成功率,同时保持了良好的安全-效用权衡。在重复的 GLM-5 运行中,SkillSonar 将 ID ASR 从 0.482 降至 0.104,OOD ASR 从 0.606 降至 0.115。进一步分析表明,其在受害模型、保留的风险家族和外部基准间具有可迁移性,且仍能抵御自适应攻击者。 ablation 研究还表明,明确的安全责任分配和技能原生表示对观察到的增益均很重要。

英文摘要

Skill-augmented agents load reusable skills as persistent runtime context, improving task performance but also giving malicious skills a durable channel for steering future actions. Such skills may leak secrets, corrupt code, bypass approvals, or stage data for exfiltration only after a concrete user task and workspace state make the unsafe action appear useful. This makes pre-install vetting insufficient and calls for runtime, task-conditioned protection. We propose Defense-as-Skill, a defense paradigm that implements the runtime guard itself as an installable, inspectable, and editable skill. Our guard, SkillSonar, runs alongside untrusted task skills and checks sensitive actions against the user's task boundary, routing each action to an allow, replan, or confirmation decision without modifying the underlying agent runtime. To study this setting, we construct SCOPE-R, a task-conditioned dataset covering 6 risk families and 21 sub-categories, with 206 attack-confirmed malicious instances and 43 benign tasks. We then improve SkillSonar on the SCOPE-R training subset using runtime guard-skill evolution, a Monte-Carlo Tree Search procedure that evolves the on-disk guard skill from feedback on the rollouts. Across Claude Code and OpenClaw, the evolved guard substantially reduces attack success while maintaining a favorable safety-utility trade-off. On repeated GLM-5 runs, SkillSonar reduces ID ASR from 0.482 to 0.104 and OOD ASR from 0.606 to 0.115. Further analyses demonstrate transfer across victim models, held-out risk families, and external benchmarks, as well as retained protection against adaptive attackers. Ablations further show that explicit safety responsibility assignment and the skill-native representation are both important to the observed gains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑