arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CompoSkill:通过可单独通过扫描器的大语言模型智能体技能实现的组合技能链攻击

CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills

Mingxiao Liu, Zhoumian Jiang, Jianan Ma, Jian Zhang, Jialuo Chen, Xinhao Deng, Zhen Wang

arXiv 2608.16246首次发表:更新:

发表机构

Hangzhou Dianzi University; Ant Group; Zhejiang University; Tsinghua University(杭州电子科技大学; 蚂蚁集团; 浙江大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CompoSkill框架揭示现有单技能认证的自主AI智能体存在系统性缺口,其通过双攻击者系统构造组合技能链攻击,在白、黑盒设置下分别实现83.3%、80.6%的风险链形成率,现有扫描器拦截效果有限。

AI 中文摘要

处理长程任务的自主AI智能体依赖于市场上逐一认证的技能:扫描器对每项技能返回安全判定,若所有包都通过则声明生态系统安全。我们表明,这一假设在技能组合下不成立:某项技能可能单独通过单个技能扫描器,但当智能体将其输出、能力或副作用与其他通过扫描器的技能的输出、能力或副作用相结合时,会参与到危险的组合中。这使得技能组合风险成为路径级属性而非节点级属性,解释了为何现有检查单个包的技能扫描器拦截能力有限。为研究这一威胁,我们提出CompoSkill框架,其通过双攻击者系统构建技能组合攻击:白盒攻击者知晓受害者已安装的技能池,直接注入明确的技能ID序列;黑盒攻击者仅知晓角色配置文件,下载对应场景的顶级市场技能,构建技能组合图,并搜索高风险链,其中隐含诱饵从不提及技能标识符。我们进一步构建CompoSkill-Bench,这是一个包含1140条记录的基准,基于OpenClaw和Nanobot上的5类威胁、6个场景的长程专业工作流构建。CompoSkill在白盒设置下实现风险链形成率(CFR)高达83.3%,在黑盒设置下为80.6%,而现有技能扫描器仅拦截了有限比例的危险组合。最后,我们观察到“桥梁增益-随后衰减”模式:桥梁技能可提高攻击成功率,但当风险链超过3个技能时,攻击成功率(ASR)会因额外跳数而下降。这些结果揭示了自主AI智能体在单技能认证方面存在系统性缺口。

英文摘要

Autonomous AI agents tackling Long Horizon Tasks depend on marketplace skills that are certified one at a time: a scanner returns a safety verdict for each skill and declares the ecosystem safe if every package passes. We show that this assumption fails under skill composition. A skill may pass the per-skill scanner individually yet participate in a risky composition when an agent connects its outputs, capabilities, or side effects with those of other scanner-passing skills. This makes skill composition risk a path level property rather than a node level property, explaining why existing skill scanners that inspect individual packages achieve limited interception. To study this threat, we present CompoSkill, a framework that constructs skill composition attacks through a dual attacker system. The white-box attacker knows the victim's installed skill pool and directly injects explicit skill-id sequences; the black-box attacker knows only a role profile, downloads the top marketplace skills for that scenario, builds a Skill Composition Graph, and searches for high risk chains whose implicit lures never name skill identifiers. We further construct CompoSkill-Bench, a benchmark of 1,140 records built from long-horizon professional workflows across five threats and six scenarios on OpenClaw and Nanobot. CompoSkill achieves risk Chain Formation Rates (CFR) up to 83.3% in the white box setting and 80.6% in the black box setting, while existing skill scanners block only a limited fraction of the risky compositions. Finally, we observe a bridge-bonus-then-hop-decay pattern: a bridge skill can increase attack success, but Attack Success Rate (ASR) decreases once additional hops make the risk chain longer than three skills. These results expose a systematic gap in single skill certification for autonomous AI agents.

Comments9 pages,5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑