发表机构
Beihang University; Baidu Inc; Wilfrid Laurier University(北京航空航天大学; 百度公司; 威尔弗里德·劳雷尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种针对LLM智能体的技能投毒攻击范式,将恶意执行与情境借口解耦,通过上游接地技能伪造理由,使下游完整执行隐藏于合法任务中,并开发自动化框架实现高成功率攻击,暴露安全审计盲点。
AI 中文摘要
LLM智能体越来越依赖可重用的技能(Skill)来处理复杂的多步骤任务,这形成了一个关键的供应链攻击面,其中被投毒的技能内容在良性请求下会操纵智能体的决策循环。现有的技能投毒攻击要么将执行行为与其情境借口放在一起,要么将执行行为分散在多个技能中,但都没有明确地将执行的理由与操作本身分开。在这项工作中,我们揭示,不受信任的智能体决策从根本上依赖于两个概念上不同的风险实现因素(RRFs):执行因素(指定执行什么具体操作)和借口因素(提供为什么智能体必须执行该操作的情境理由)。在此抽象指导下,我们提出了一种基于协调的攻击范式:将借口与执行解耦。我们不将恶意执行行为碎片化,而是将其作为一个完整的操作保留在下游的引导技能(Steering Skill)中,同时将借口因素委托给上游的接地技能(Grounding Skill),该技能通过常规的实用操作微妙地改变持久的环境工件。因此,完整的执行行为就隐藏在众目睽睽之下,只有在对照捏造的借口进行评估时,它才显得完全合法且由任务驱动。基于这一表述,我们开发了一个自动化框架,该框架发现真实的执行依赖关系,合成协调的借口-执行技能对,并通过运行时闭环反馈迭代优化被投毒的技能指令。跨单会话和持久跨生命周期场景的广泛评估表明,解耦技能投毒实现了高攻击成功率,暴露了孤立技能安全审计中的一个关键盲点。我们的自动化框架代码可在以下网址获取:https://this-url。
英文摘要
LLM agents increasingly rely on reusable Skills for complex, multi-step tasks, creating a critical supply-chain attack surface where poisoned Skill content steers agent decision loops under benign requests. Existing skill poisoning attacks either colocate actuation with its contextual pretext or distribute actuation across multiple Skills, but do not explicitly separate the rationale for execution from the operation itself. In this work, we reveal that untrusted agent decisions fundamentally depend on two conceptually distinct Risk-Realization Factors (RRFs): an actuation factor (specifying what concrete operation is performed) and a pretext factor (providing the situational rationale for why the agent must perform it). Guided by this abstraction, we propose a coordination-based attack paradigm: decoupling pretext from actuation. Rather than fragmenting the malicious actuation, we preserve it as an intact operation within a downstream Steering Skill, while delegating the pretext factor to an upstream Grounding Skill that subtly alters persistent environment artifacts through routine utility operations. The intact actuation thus hides in plain sight, appearing completely legitimate and task-driven only when evaluated against the fabricated pretext. Building on this formulation, we develop an automated framework that discovers authentic execution dependencies, synthesizes coordinated pretext-actuation skill pairs, and iteratively refines poisoned skill instructions via runtime closed-loop feedback. Extensive evaluations across single-session and persistent cross-lifecycle scenarios demonstrate that decoupled skill poisoning achieves high attack success, exposing a critical blind spot in isolated Skill security audits. Our automated framework code is available at https://github.com/Wenxin-buaa/CoordPoison.git.