发表机构
Max Planck Institute for Intelligent Systems; Tel Aviv University; Sapienza University of Rome; University of Cagliari; ELLIS Institute Tübingen(马克斯·普朗克智能系统研究所; 特拉维夫大学; 罗马第一大学; 卡利亚里大学; 图宾根ELLIS研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多技能环境中现有技能注入攻击成功率被高估的问题,提出路由器感知攻击CORSA,通过聚类优化同时提升检索与执行成功率,实验显示显著优于现有方法且保持用户效用。
AI 中文摘要
AI智能体日益依赖模块化的第三方“技能”,这些技能由技能路由器动态选择以执行复杂任务。尽管近期研究强调了嵌入这些技能中的提示注入威胁,但现有评估通常假设恶意技能已被选中执行。我们表明,这一假设可能大幅高估攻击成功率。在现实的多技能环境中,注入的技能必须首先竞争检索,这使得现有注入的有效攻击成功率(ASR)降低了87-97%。为解决这一局限,我们引入了CORSA(面向路由器感知技能攻击的聚类优化),一种路由器感知攻击,它跨相关任务聚类优化技能注入以实现检索与执行。我们通过扩展SkillRouter引入的基准,增加八类恶意载荷,评估了路由器管理下多技能环境中的技能注入攻击。CORSA采用连续优化阶段,首先改进检索,然后优化端到端攻击成功率,同时我们分别评估用户效用和注入自然度。我们的实验表明,与现有技能注入相比,CORSA在保持用户效用的同时,显著提高了检索和端到端攻击成功率,且所得攻击可跨不同路由器架构和LLM骨干迁移。
英文摘要
AI agents increasingly rely on modular third-party "skills" that are dynamically selected by skill routers to execute complex tasks. While recent studies highlight the threat of prompt injections embedded in these skills, existing evaluations often assume settings where the malicious skill is already selected for execution. We show that this assumption can substantially overestimate attack success. In realistic multi-skill environments, injected skills must first compete for retrieval, reducing the effective attack success rate (ASR) of existing injections by 87-97%. To address this limitation, we introduce CORSA (Cluster Optimization for Router-Aware Skill Attacks), a router-aware attack that optimizes skill injections for both retrieval and execution across clusters of related tasks. We evaluate skill injection attacks under router-managed multi-skill settings by extending the benchmark introduced by SkillRouter with eight malicious payload categories. CORSA uses successive optimization stages to first improve retrieval and then optimize end-to-end attack success, while we evaluate user utility and injection naturalism separately. Our experiments show that CORSA substantially improves both retrieval and end-to-end attack success over existing skill injections while preserving user utility, and that the resulting attacks transfer across different router architectures and LLM backbones.
Comments16 pages, 5 figures