AI 中文总结
该研究提出SkillShift框架,可隐蔽引导LLM智能体策略,在保持高效用的同时实现攻击者偏好选择率,且现有扫描器难以检测,需对智能体技能开展行为审计。
AI 中文摘要
可复用智能体技能为大语言模型(LLM)智能体提供任务流程、工具使用指导和输出约束,但这些技能也作为外部化行为策略,产生供应链风险:第三方技能可能在保留声明任务和有效输出接口的同时,隐蔽地将智能体决策重定向至未公开目标。我们将技能策略完整性形式化,要求技能诱导的策略与其声明功能和用户授权目标保持一致。我们进一步提出SkillShift,这是一种用于隐蔽策略引导的受限黑盒框架,无需显式目标命令注入或任务劫持。它结合语义合理的策略编辑与分层验证、失败引导优化和策略压缩,以保持有效性、输出有效性、可迁移性和隐蔽性。我们在智能体商业和软件依赖使用中实例化此威胁,SkillShift在维持100%效用保留率的同时,实现了攻击者偏好的81.33%和63.33%的选择率。冻结的策略无需进一步优化即可跨异构LLM后端和智能体环境迁移。此外,被评估的扫描器无法检测到构建的技能,这促使将可复用技能作为智能体策略工件进行行为审计。
英文摘要
Reusable agent skills extend large language model (LLM) agents with task procedures, tool-use guidance, and output constraints. Yet these skills also act as externalized behavioral policies, which create a supply-chain risk: a third-party skill may preserve the declared task and valid output interface while covertly redirecting agent decisions toward an undisclosed objective. We formalize Skill Policy Integrity, which requires a Skill-induced policy to remain aligned with its declared functionality and the user-authorized objective. We further present SkillShift, a constrained black-box framework for covert policy steering without explicit target command injection or task hijacking. It combines semantically plausible policy edits with hierarchical validation, failure-guided optimization, and strategy compression to preserve effectiveness, output validity, transferability, and inconspicuousness. We instantiate this threat in agentic commerce and software dependency use, with SkillShift achieving attacker-favored selection rates of 81.33% and 63.33% while maintaining a 100% utility-preserving rate. The frozen policies also transfer without further optimization across heterogeneous LLM backends and agent environments. Moreover, the evaluated scanners fail to detect the constructed skills, motivating behavioral auditing of reusable skills as agent policy artifacts.
Comments25 pages