发表机构
Huawei Research(华为研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出白盒LLM攻击者Pretext,通过将恶意载荷移至自然语言并拆分指令,成功逃避AI智能体技能检测框架,最高逃避率达97%,揭示现有扫描器漏洞。
AI 中文摘要
技能通过向上下文中注入指令和信息来扩展智能体的能力,并被OpenClaw和Claude Code等智能体广泛使用。先前的研究表明,第三方市场托管着恶意技能,这些技能使攻击者能够直接影响受害者的智能体。新兴的防御措施在安装前扫描技能,将确定性静态检查与基于LLM的语义判断器配对,如NVIDIA的SkillSpector。我们证明,此类防御会被了解检测器的攻击者所击败。我们的白盒LLM攻击者Pretext迭代地制作技能,既能逃避检测,同时仍能传递载荷并执行良性任务:将载荷从代码转移到自然语言中使静态分析失效,而将其伪装为技能的合法目的并在文件间拆分指令,使LLM阶段低于其阻断阈值。在三个开源模型上,Pretext针对冻结检测器实现了高达97%的逃避率,针对共同适应检测器实现了77%的逃避率,揭示了当前技能扫描器中的重大漏洞。
英文摘要
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97\% and 77\% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.
CommentsAccepted in AIWild@NeurIPS 2026