AI 中文总结
研究针对编码代理中恶意技能文件检测问题,提出SkillGate安全网关,采用混合正则表达式预过滤器+大语言模型判断管道,在SkillsBench基准测试中,检测效果好、成本低、运行时开销小,相比现有工具性能显著提升。
AI 中文摘要
软件工程团队现在将人工智能编码代理(Cursor、Claude Code、GitHub Copilot)作为一流的生产力工具,安装特定领域的技能文件以根据项目 API、框架约定和组织工作流程调整代理行为。这些复杂的Markdown文件通过简单命令就能从公共注册表轻松下载且无安全筛选,存在供应链攻击风险,恶意技能文件可悄悄改变代理行为。目前尚无针对此攻击面的系统工具链防御。我们提出了SkillGate,一个在编码代理安装前筛选人工智能技能包的可部署安全网关。SkillGate使用混合正则表达式预过滤器+大语言模型判断管道,能节省成本。我们针对两个现有工具回答了四个研究问题,在SkillsBench基准测试中,SkillGate实现了F1=0.817,FPR=1.13%,同时将大语言模型输入令牌减少了77%,在与阈值无关的AUPRC上比现有工具性能高出5 - 6倍。
英文摘要
Software engineering teams now deploy AI coding agents (Cursor, Claude Code, GitHub Copilot) as first-class productivity tools, installing domain-specific skill files to tailor agent behavior to project APIs, framework conventions, and organizational workflows. These complex Markdown files are easily downloaded from public registries with a single npx skills add command and no real security screening, representing a novel supply-chain attack surface: a malicious skill file can silently reprogram agent behavior, exfiltrating credentials, injecting backdoors into generated code, or redirecting agent actions to attacker-controlled endpoints. The threat is not hypothetical: recent reports document hundreds of malicious skill packages in public registries, including organized campaigns that distributed credential-stealing infostealers via fake productivity skills. No systematic toolchain defense exists for this attack surface. We present SkillGate, a deployable security gateway that screens AI skill packages before coding agent installation. SkillGate uses a hybrid regex-prefilter + LLM-judge pipeline: safe-signal files bypass the LLM entirely (skip savings); flagged files have only their matched snippet windows sent to the judge, not the full content (snippet savings). We answer four research questions covering detection effectiveness, screening cost, runtime overhead, and false positive behavior on the SkillsBench benchmark against two existing tools. On SkillsBench (n=1,650, 9.1% malicious), SkillGate achieves F1=0.817, FPR=1.13% while reducing LLM input tokens by 77% vs. full-file screening, and outperforming existing tools by 5-6x on threshold-independent AUPRC (0.830 vs. 0.144/0.162).
Comments10 pages, 5 figures