arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29744cs.SEcs.AI

提交之间:全AI编写代码库中的过程、错误与声明可靠性

Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase

发表机构都柏林圣三一学院
查看机构详情
  • Trinity College Dublin(都柏林圣三一学院)

机构由 AI 辅助整理,请以论文原文为准。

Douglas Leith

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过分析一个完全由AI编写的21,000行Python代码库,提出数据集、溯源工具和分类体系,发现AI代码开发主动性强,但存在14.3%的真实错误率和约20-25%的事实错误响应。

中文摘要 AI 辅助

我们提出了:(i)一个新数据集,包含一个完全由Claude AI构建的21,000行Python工具的完整开发历史,其中没有人类编写的代码或测试;(ii)两个代码溯源追踪工具;(iii)三个分别针对指令意图、提交溯源和响应可靠性的分类体系;(iv)将这些工具和分类体系应用于分析该数据集。我们发现:(i)用户编码代理的CLI指令与IDE聊天指令在类型上有所不同,更侧重于理解、规划和咨询;(ii)代码开发主要是主动式的;(iii)14.3%的AI代码生成事件包含一个真实错误,该错误后来被AI编写的测试套件捕获;(iv)AI的交互式响应中大约有四分之一到五分之一包含一个或多个事实错误。

英文摘要

We present: (i) a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI, with no human-authored code or tests, (ii) two code-provenance tracing tools, (iii) three taxonomies for instruction intent, commit provenance, and response reliability, (iv) application of these to analyse the dataset. We find that: (i) user coding agent CLI instructions differ in kind from IDE-chat instructions, with a greater focus on comprehension, planning and consultation, (ii) code development is mainly proactive, (iii) 14.3% of AI code-generation events contain a real error later caught by the AI-authored test suite, (iv) roughly 1 in 4-5 of the AI's interactive responses contains one or more factual errors.

↑