单独隐蔽,协同危害:基于技能智能体系统中的技能级联攻击
Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems
浏览论文内容
中文总结 AI 辅助
提出技能级联攻击范式,揭示多技能协同可规避检测并造成危害,开发SkillCascade框架与基准,强调需跨技能交互防御。
中文摘要 AI 辅助
技能是一种模块化包,包含自然语言指令、可执行脚本和参考资源,智能体可在运行时加载以扩展其针对特定任务的能力。基于技能的智能体系统因此能够灵活复用第三方能力,但这种技能生态系统的开放性也开辟了新的攻击面。先前的工作主要关注单个技能内部的漏洞,但很少关注跨技能交互所产生的风险。在本文中,我们引入了技能级联攻击,这是一种威胁范式,其中恶意目标被分布在多个技能中,使得每个修改在孤立状态下看似良性,但它们的组合执行却是有害的。例如,在处方审查流程中,第一个技能削弱了提取病史中最近停用药物信号,第二个技能降低了与之相关的任何药物相互作用的严重性,第三个技能在最终摘要中抑制了由此产生的低优先级警报,从而使得严重的药物相互作用警告在到达医生之前悄无声息地消失。为了系统地研究这一安全盲点,我们开发了SkillCascade,一个自动化的多智能体红队框架,并发布了SkillCascade-Bench,一个包含213个经过验证的级联测试用例的基准,涵盖多个智能体系统和领域。在代表性智能体(如OpenClaw、Claude Code、Codex)和LLM骨干上,级联交互可靠地诱导有害行为,同时规避现有的逐技能扫描器和运行时监控器。我们的研究结果凸显了组件级完整性与系统级安全性之间的差距,并呼吁防御措施应推理跨技能交互,而非孤立地对待单个技能。
英文摘要
A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface. Prior work has focused on vulnerabilities within individual skills, but little attention has been paid to risks that arise from interactions across skills. In this paper, we introduce skill cascading attacks, a threat paradigm in which a malicious objective is distributed across multiple skills so that each modification looks benign in isolation, yet their combined execution is harmful. For instance, in a prescription-review pipeline, the first skill weakens signals of recently discontinued medications in the extracted history, the second downgrades the severity of any drug interaction tied to them, and the third suppresses the resulting low-priority alert in the final summary, so that a severe drug-interaction warning silently disappears before reaching the physician. To systematically study this safety blind spot, we develop SkillCascade, an automated multi-agent red-teaming framework, and release SkillCascade-Bench, a benchmark of 213 validated cascading test cases across multiple agent systems and domains. Across representative agents (e.g., OpenClaw, Claude Code, Codex) and LLM backbones, cascaded interactions reliably induce harmful behaviors while evading existing per-skill scanners and runtime monitors. Our findings highlight a gap between component-level integrity and system-level safety, and call for defenses that reason over cross-skill interactions rather than individual skills in isolation.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- University at Buffalo, SUNY(纽约州立大学布法罗分校)
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。