arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扫描缰绳:AI 编码智能体配置中供应链缺陷的实证研究

Scanning the Harness: Configuration Exposures in AI Coding-Agent Supply Chains

Benjamin Kapner, Carmel Soceanu, Alicia Petrunin, Hofni Gartner

arXiv 2609.07360首次发表:更新:

发表机构

Red Hat; Stein Faculty of Computer and Information Science, Ben-Gurion University of the Negev(红帽; 内盖夫本古里安大学斯坦计算机与信息科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究实证分析 AI 编码智能体配置供应链,发现 16.0% 的设置存在安全缺陷,并提出基于字节的验证方法,强调配置依赖风险。

AI 中文摘要

AI 编码智能体(如 Claude Code、Cursor、GitHub Copilot 和 OpenAI Codex)通过开发者编写和共享的工件进行配置:指令文件、技能、钩子、MCP 服务器声明、子智能体。这个缰绳是一个从市场和公共仓库安装的依赖层,以开发者的权限运行,没有锁文件,没有安装时检查,也没有描述组件可能做什么的词汇表。我们对其进行了研究,涉及 3,171 个公共 GitHub 仓库:2,660 个组装了两种或更多组件类型的设置和 511 个已发布的技能集合。我们仅测量可从字节判定且后果为安全暴露、无法工作的配置或偏离 Agent Skills 规范的规则,并在每个发现计入之前进行验证:一个独立的实现从其固定提交的仓库中重新推导出该发现,一个带有已发布提示的语言模型裁决者裁决每个分歧,第二个独立的模型会话重新检查每个计入的对。三类安全问题得以确认:9.8% 的设置安装了未固定版本的 MCP 服务器,3.1% 在看似范围受限的授权(如 Bash(python:*))背后预先批准任意执行,3.8% 携带预先批准 shell 的技能,无论谁安装它。总计 16.0% 的设置携带安全缺陷,16.7% 携带任何类型的已确认缺陷,而相同规则下的原始扫描器率为 25.5%;第三类问题出现在 3.7% 的集合中,市场扫描可以看到它。比较两个文件的规则检测到通常是有意的差异,并作为观察结果报告。未确认任何凭据泄露路径。该工具、语料库清单、提示和每个裁决均已发布。

英文摘要

AI coding agents rely on repository instructions, skills, hooks, tool-server declarations, and subagent definitions. These artifacts distribute both behavior and access to executable dependencies, making configuration review part of the agent software supply chain. We study 3,171 public GitHub repositories: 2,660 assembled setups and 511 skill collections. Deterministic analysis, mechanical re-derivation, model-assisted adjudication, human review, and platform documentation checks identify six categories of configuration exposure and conformance issues. Unpinned MCP package declarations occur in 9.8% of setups, broad execution grants in 2.5%, and broad skill tool preapproval in 3.8%. Their union covers 409 setups (15.4%); among setups with MCP configuration, 24.5% contain an unpinned declaration. Including required-field and skill-format issues brings the setup rate to 17.9% and the collection rate to 6.8%. Holding the six categories fixed, contextual review changes the setup rate from 18.3% to 17.9%; rule selection explains most of the reduction from the broader candidate set. The findings identify concrete opportunities to pin dependencies, review execution pre approval, and check component conformance before distribution or use. A documented permission exception also exposes a shared error in the scanner and its mechanical audit, motivating version-specific semantic checks. The study provides reproducible evidence about repository declarations; agreement with human judgments informs label validation, while runtime consequences and recall remain unmeasured.

Comments8 pages, 1 figure, 5 tables. Revised exposure definitions, corrected counts, and clarified methodology. Tool: https://github.com/redhat-community-ai-tools/harness-eval; data and scripts: https://github.com/Benkapner/harness-eval-experiments

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑