发表机构
Nankai University; Northeastern University; University of Science and Technology Beijing; Southeast University; Nanyang Technological University(南开大学; 东北大学; 北京科技大学; 东南大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出ColluSkill攻击框架,可通过跨技能组合规避技能扫描器,攻击成功率达96.0%;还提出ChainGuard防御框架,能将攻击成功率降至22.5%,同时放行99.5%良性工作流。
AI 中文摘要
智能体技能正成为基于大语言模型(LLM)的智能体系统中重要的攻击面。通过对现有技能扫描器的实证研究,我们发现当前防御措施主要检查单个技能,对跨技能组合带来的风险研究不足,这形成了一个实际盲点:多个局部看似合理的技能可能通过安全检查,但在智能体执行时会共同构成有害工作流。为探究该威胁,我们提出ColluSkill,这是一种合谋式多技能链攻击框架,它将完整恶意意图分解为嵌入独立打包技能中的相互依赖的子有效载荷。该攻击不依赖任何单个恶意技能,而是通过上下文依赖、工件传递和执行交接,从局部合理行为的有序组合中产生。ColluSkill还采用基于LLM的链规划和扫描器反馈优化,以保留链级攻击语义,同时减少单个子技能中的可疑信号。为防御此类攻击,我们提出ChainGuard,这是一种上下文感知的技能链扫描器,可联合分析候选技能与智能体环境中已安装的技能,重构跨技能依赖、工件流、能力组合和下游行为,以识别仅在工作流层面出现的风险。在6个代表性技能扫描器上的实验表明,ColluSkill的平均攻击成功率达96.0%,始终优于所评估的单技能和多技能攻击基线;同时,ChainGuard将攻击成功率降至22.5%,同时允许99.5%的良性工作流通过,凸显了智能体技能生态系统中链级安全分析的重要性。
英文摘要
Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners, we find that current defenses mainly inspect individual skills, leaving risks from cross-skill composition insufficiently examined. This creates a practical blind spot: multiple locally plausible skills may pass security checks while collectively forming a harmful workflow during agent execution. To investigate this threat, we propose ColluSkill, a collusive multi-skill-chain attack framework that decomposes a complete malicious intent into interdependent sub-payloads embedded in independently packaged skills. The attack does not rely on any single malicious skill, but emerges from the ordered composition of locally plausible behaviors through contextual dependencies, artifact passing, and execution handoffs. ColluSkill further employs LLM-based chain planning and scanner-feedback refinement to preserve chain-level attack semantics while reducing suspicious signals in individual sub-skills. To defend against such attacks, we propose ChainGuard, a context-aware skill-chain scanner that jointly analyzes a candidate skill and the skills already installed in the agent environment. ChainGuard reconstructs cross-skill dependencies, artifact flows, capability compositions, and downstream behaviors to identify risks that emerge only at the workflow level. Experiments on six representative skill scanners show that ColluSkill achieves an average attack success rate of 96.0% and consistently outperforms the evaluated single-skill and multi-skill attack baselines. Meanwhile, ChainGuard reduces the attack success rate to 22.5% while allowing 99.5% of benign workflows to pass, highlighting the importance of chain-level security analysis for agent skill ecosystems.
Comments9 pages, 3 figures, 4 tables