AI 中文总结
该研究分析智能体软件开发中责任归属的矛盾,替换验证分类法,发现责任归属方向相反,指出智能体工具未造成责任差距,相关数据已公开。
AI 中文摘要
编码智能体会在生产仓库中提交代码、发起拉取请求(pull request)并推送代码。责任归属由两个互不关联的地方确定:平台控制智能体可执行的操作,而提供方条款则为智能体生成的内容分配责任。我们结合产生制品的工作流事件,对来自7家提供方的4款智能体编码工具及18份治理政策文件展开研究,记录每个事件中谁拥有权限、谁以何种身份执行操作、谁必须进行验证、谁承担后果,以及留存了何种制品。这些责任层级存在分歧:某家提供方禁止分配任务的开发者批准对应的拉取请求;另一家则规定,智能体可批准低于配置风险阈值的拉取请求,且能驳回评审。因此,我们将常用的强制、建议及缺失验证的三分法,替换为区分机制是否强制检查与谁执行检查的网格。不同提供方的责任归属方向相反,且未定义智能体作者身份的尾注(trailer),不过某家提供方将合著者尾注(co-authorship trailer)改作此用途。我们认为,智能体工具并未造成这一差距:十年的代码审查研究已记录到,批准制品承载的责任少于条款所假定的内容。变化在于,这一弱点从人类工作方式的属性转变为产品的属性:供应商如今记录到,某产品会出现在批准事件中,且在无任何能形成判断的一方在场的情况下生成相同制品。我们并非声称该差距会损害任何人;4款工具的选择规则是一致的,每一次报告的缺失都针对翻倍的页面集进行了重新测试,并报告了留存率,且源集合及其脚本已被存入。
英文摘要
Coding agents author commits, open pull requests and push code in production repositories. Responsibility is settled in two layers that do not refer to each other: platform controls gating what an agent may do, and provider terms allocating responsibility for its output. Objective: Where the two layers disagree at a workflow event, and what each states there about authority, execution, verification, consequence and record. Method: A qualitative document study of 121 items archived byte-exact: documentation for four agentic coding tools, the platform controls on their output, and eighteen policy documents from seven providers, read by deductive content analysis against an a-priori system of five dimensions and nine workflow events. An independent second coder blind-recoded the verification-mechanism classification, a complete enumeration and not a sample (Cohen's kappa 0.81, n = 13, on whether a mechanism compels; 0.75, n = 12, on who performs it). Results: Cursor's terms make the user responsible for evaluating the use of any suggestion; its documentation describes a product that performs that evaluation and records it as an approval. No collected artifact records that the verification the terms make the user's duty took place. At merge two tools compel a person, one documents an agent approving below a configured risk threshold, and one only advises. Six of the fifteen mechanisms that compel a check have a default the vendor states; nine have one this study inferred. Readers who had not made them refuted five of the eight absence claims; three survived, one materially qualified. Conclusions: The terms attach duty and consequence to output as a class, the platform records events, nothing records the duty discharged. One vendor documents a product that forms the approval judgement, and the terms allocate consequence against the artifact regardless.
Comments30 pages, 5 tables. Source collection of 118 archived documents and 12 processing scripts deposited at Zenodo, doi:10.5281/zenodo.21965182