发表机构
CUHK-Shenzhen; SLAI(香港中文大学(深圳); 深圳市人工智能与机器人研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SkillBloat两阶段框架,通过技能注入对LLM编码智能体实施令牌放大攻击,在真实基准上实现5.4184x-10.1455x平均最佳放大,揭示了技能生态系统的新型资源攻击面。
AI 中文摘要
智能体技能为编码智能体提供了特定任务的指令、脚本和资源,同时也形成了可被滥用的可信指令通道,超出了传统安全攻击的范畴。本文研究通过技能注入实现的令牌放大攻击,这是一种经济资源滥用威胁,恶意技能会导致智能体消耗远超正常任务执行所需的令牌。我们提出SkillBloat,这是一个两阶段框架:首先在多种放大机制下筛选不同攻击类型条件的库,再通过大语言模型引导的全文档技能重写优化最强候选方案。在真实世界的技能基准上评估,SkillBloat在多个编码智能体目标配置下实现了5.4184倍至10.1455倍的平均最佳放大效果。 ablation研究显示,第二阶段优化循环相比仅第一阶段的攻击类型筛选,始终能提升平均最佳放大效果,表明迭代优化比初始攻击类型选择带来了额外收益。这些结果表明,技能生态系统暴露了与现有面向安全的技能投毒正交的实用资源放大攻击面。
英文摘要
Agent skills extend coding agents with task-specific instructions, scripts, and resources, but they also create a trusted instruction channel that can be abused beyond conventional security attacks. This paper studies token amplification through skill injection: an economic resource-abuse threat in which a malicious skill causes an agent to consume substantially more tokens than needed for normal task execution. We present SkillBloat, a two-phase framework that first screens a library of diverse attack-type conditions across multiple amplification mechanisms and then refines the strongest candidate through LLM-guided full-document skill rewriting. Evaluated on a real-world skill benchmark, SkillBloat achieves 5.4184x-10.1455x average best amplification across multiple coding-agent target configurations. An ablation shows that the second-stage refinement loop consistently improves average best amplification over Phase 1 attack-type screening alone, demonstrating that iterative optimization provides additional benefit beyond initial attack-type selection. These results show that skill ecosystems expose a practical resource-amplification attack surface that is orthogonal to existing security-oriented skill poisoning.