arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillSecurer:检测与修复AI智能体技能中的提示注入漏洞

SkillSecurer: Detecting and Patching Prompt-Injection Vulnerabilities in AI Agent Skills

Donato Mecca, Alberto Verna, Youness Bouchari, Nikhil Jha, Marco Mellia

arXiv 2609.14079首次发表:更新:

发表机构

Politecnico di Torino(都灵理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体技能易受提示注入攻击的问题,提出全智能体框架SkillSecurer,通过红蓝智能体生成、检测、定位并修复漏洞,实现100%注入检测率,并发现超17%热门技能存在潜在风险。

AI 中文摘要

智能体技能通过可复用的指令、脚本和配置扩展了AI智能体的能力,但同时也面临着新的攻击方式,这些攻击可能影响智能体的决策和行动。为应对这些风险,我们提出了SkillSecurer,一个完全智能体化的框架,用于生成、检测、定位并修复智能体技能中的安全风险。其红色智能体能够跨九种威胁类型生成与上下文兼容的注入,并记录精确的修改;其蓝色智能体则分析完整的技能包,生成有根据的证据,并提出补丁。对于受控实例,验证器将发现结果和补丁与记录的注入进行比较,从而实现注入级别的评估。我们通过选择最佳后端大语言模型、与竞争对手进行比较,并手动交叉验证每个评估阶段,对SkillSecurer进行了全面评估。在其最佳性能后端的支持下,SkillSecurer是唯一实现100%注入检测率的扫描器。接下来,我们分析了来自此HTTP URL的热门技能,发现所检查的技能中有超过17%存在潜在漏洞。通过测试其中一些技能,我们触发了实际事件,展示了运行未经验证技能的风险。我们的结果表明,上下文感知的大语言模型分析能够提供可靠的注入定位和可操作的修复,而不仅仅是技能级别的标记。

英文摘要

Agent skills extend AI agents with reusable instructions, scripts, and configuration, but are also open to new attacks to influence an agent's decisions and actions. To address these risks, we present SkillSecurer, a fully agentic framework for generating, detecting, localising, and remediating security risks in agent skills. Its red agent generates context-compatible injections across nine threat types while recording the exact modification; its blue agent analyses complete skill packages, produces grounded evidence, and proposes patches. For controlled instances, a verifier compares findings and patches with the recorded injection, enabling injection-level evaluation. We thoroughly evaluate SkillSecurer by selecting the best backend LLM, comparing it with competitors, and manually cross-validating each evaluation stage. With its best performing backend, SkillSecurer is the only scanner to achieve a 100% injection detection rate. Next, we analyse popular skills from skills.sh, finding latent vulnerabilities in more than 17% of the skills examined. Testing some of those skills, we trigger actual incidents, showing the risks of running unverified skills. Our results show that context-aware LLM analysis can provide reliable injection localisation and actionable remediation beyond skill-level flagging alone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑