arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 编程助手在安装前会检查吗?研究软件供应链中信任信号的需求方审计预注册研究

Do AI Coding Assistants Check Before They Install? A Pre-Registered Demand-Side Audit of Trust Signals in the Research Software Supply Chain

Pengyin Shan

arXiv 2609.07754首次发表:更新:

发表机构

National Center for Supercomputing Applications; University of Illinois Urbana-Champaign(国家超级计算应用中心; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过预注册审计,发现AI编程助手在安装研究软件时极少检查信任信号,验证率仅0.5%,且价格与验证能力无关,提出应将验证内置于助手程序中。

AI 中文摘要

AI 编程助手现在能够选择、安装和配置软件,而攻击者已通过虚构包名、入侵维护者账户和操纵仓库文本利用了这一地位。为此,供应链社区发布了机器可检查的信任信号:软件物料清单、签名发布、构建来源证明和声明的官方渠道。对于研究软件,这些信号类别中,编程助手是否读取或依据这些信号行动尚未被测量。我们预注册并开展了一项受控研究,涉及从 87 个项目语料库中选取的六个开源研究软件项目(三个高性能计算,三个量子计算),并在任何试验前将协议、种子、面板和分析计划以 DOI 形式存放。我们为每个项目创建了九个修改副本:无信号、每类信号各一个、两个带有错误签发者的签名或证明、一个包含全部四种信号、以及一个复现项目自身元数据中记录的冲突。三种模型在两种操作助手方式下(有审批步骤和无审批步骤)进行了 1,920 次注册试验,另加三个前沿模型的补充试验。我们根据容器日志而非助手所述来评分行为,并记录了每次试验的成本。在所有条件下,验证都很少发生:在 1,920 次注册试验中,有 9 次(0.5%)助手在安装前打开了任何来源信号;在 384 次对照试验中,有 0 次;且没有试验运行验证命令,因此信号的存在没有可测量的效果。我们得出三个结论:发布信号是必要的但不充分;价格并未带来验证(验证最多的模型每次试验成本为 0.10 美元;最强大的模型成本为 1.00 美元,却未验证任何内容);验证必须内置于运行助手的程序中。我们发布了每次试验的成本账本、协议和所有日志。

英文摘要

AI coding assistants now select, install, and configure software, and attackers have exploited that position through invented package names, compromised maintainer accounts, and manipulated repository text. In response, the supply-chain community publishes machine-checkable trust signals: software bills of materials, signed releases, build provenance attestations, and declared official channels. Whether coding assistants read or act on those signals has not been measured for any of these classes on research software. We pre-registered and ran a controlled study on six open-source research software projects (three HPC, three quantum computing) drawn from an 87-project corpus, with protocol, seed, panel, and analysis plan deposited with a DOI before any trial. W created nine modified copies for each project: no signal, one per signal class, two with a signature or attestation from the wrong issuer, one with all four signals, and one reproducing documented conflicts in the project's own metadata. Three models under two ways of operating an assistant, with and without an approval step, gave 1,920 registered trials, plus a supplement on three frontier models. We scored behavior from container logs rather than from what the assistant said, and recorded the cost of every trial. Verification was rare under every condition: in 9 of 1,920 registered trials (0.5%), the assistant opened any provenance signal before installing in 0 of 384 control trials, and no trial ran a verification command, so signal presence had no measurable effect. We drew three conclusions: publishing signals is necessary but not sufficient; price did not buy verification (the model that verified most often costs $0.10 per trial; the most capable, at $1.00, verified nothing); verification must be built into the program that runs the assistant. We release the per-trial cost ledger, the protocol, and every log.

Comments18 pages, 11 figures, 8 tables. Pre-registered protocol deposited 2026-08-22 (10.5281/zenodo.22062503); tooling v0.2.1 at 10.5281/zenodo.22544144; data at 10.5281/zenodo.22546062

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑