发表机构
The University of Hong Kong; Shandong University; Shanghai Jiao Tong University; Southeast University(香港大学; 山东大学; 上海交通大学; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出APEX方法,通过构建对抗性技能链,利用技能间信息传递携带虚假用户批准声明,诱导LLM智能体执行攻击者选定操作,在SkillsBench上成功率高达74.2%,并揭示现有防御的局限性。
AI 中文摘要
LLM智能体利用技能来提升在专门任务上的表现。为了完成用户请求,智能体可能按顺序调用多个技能,使得在一个技能下产生的信息能够指导下一个技能。由于技能可能来自开源仓库,这种交接也可能将攻击者控制的声明带入后续决策中。在本文中,我们提出了APEX,它构建并优化针对用户任务和攻击者选定操作的对抗性技能链。关键洞察在于,智能体编写的真实任务进展记录可以携带虚假的用户批准声明跨技能传播:上游技能诱导智能体创建该记录,下游技能利用该记录来引导攻击者选定的操作。在SkillsBench上,针对四个目标操作家族和六个模型,这些链在690次尝试中成功诱导了选定操作512次(74.2%)。在GPT-5.4上,完整链在84.3%的尝试中成功,而将工作流合并为单个技能时成功率仅为17.4%。我们进一步评估了一种提示防御,要求智能体将技能生成的文件与原始请求进行核对。在GPT-5.4上,该防御将目标操作成功率从84.3%降至59.1%,而在72个良性原生技能任务中,验证器测试通过率从86.7%降至56.3%。这些结果凸显了需要既能防止攻击者导向操作又能保持合法任务性能的防御措施。
英文摘要
LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next. Because skills may come from open-source repositories, this handoff can also carry attacker-controlled claims into later decisions. In this paper, we introduce APEX, which constructs and refines adversarial skill chains tailored to a user task and an attacker-selected action. The key insight is that an agent-written record of genuine task progress can carry a false claim of user approval across skills: an upstream skill induces the agent to create the record, and a downstream skill uses it to direct the attacker-selected action. Across four targeted-action families and six models on SkillsBench, the chains induce the selected action in 512 of 690 attempts (74.2%). On GPT-5.4, the full chain succeeds in 84.3% of attempts, compared with 17.4% when the workflow is merged into one skill. We further evaluate a prompting defense that asks the agent to check skill-produced files against the original request. On GPT-5.4, it lowers targeted-action success from 84.3% to 59.1%, while the verifier test-pass rate across 72 benign native-skill tasks falls from 86.7% to 56.3%. These results highlight the need for defenses that prevent attacker-directed actions while preserving legitimate task performance.