arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从智能体行为到智能体友好型文档:关于编码智能体如何发现、阅读和编写技术文档的实证研究

From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation

Zhijun Gao, Jing Chen

arXiv 2608.20195首次发表:更新:

发表机构

Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过SWE-chat和AIDev两个数据集,揭示编码智能体的文档交互规律,发现面向智能体的制品占主导,文档与代码编辑关联弱,验证序列缺失,文档滞后代码,且"智能体友好型"文档的可操作性、可验证性缺乏行为支持。

AI 中文摘要

技术文档是为人类开发者编写的,但如今越来越多的软件变更由自主编码智能体完成。它们会查阅哪些文档、何时查阅以及后续行为仍是未知的。我们基于智能体行为,在两个公开数据集上开展智能体与文档交互的研究:SWE-chat的557次智能体编码会话,产生94813个开发事件,其中包含3033次文档交互;AIDev的33097个智能体拉取请求,包含690260条分类文件级变更记录。四项发现对当前文档实践提出挑战:第一,智能体的文档工作以面向智能体的制品为主:指令文件和工作笔记占所有文档交互的60.5%,而传统技术文档占10.6%,API参考占1.3%;第二,查阅与代码编辑的关联尚不明确:相邻转移概率为0.002,未调整的三事件提升值为1.05,而经阶段调整的模型显示其比值比(OR)大于1(OR 1.33 [1.09, 1.62]);文档创建的未调整提升值为1.67,但调整后的区间包含1;第三,未观察到明确的基于文档的验证序列,且查阅与较少的即时测试相关(提升值0.23,聚类置信区间0.08-0.45;调整后OR 0.39 [0.25, 0.60]);第四,查阅多为主动发起(70.2%),远多于失败驱动型(7.5%),且文档滞后于代码:在同时变更二者的多提交拉取请求中,代码先被触及的频率是文档的4.7倍。我们从这些痕迹中得出智能体与文档交互的描述性模型,为双叶循环而非线性流程,并表明"智能体友好型"文档的两个普遍假设属性——可操作性和可验证性——缺乏一致的行为支持。我们发布了我们的流程、编码方案和事件级数据。

英文摘要

Technical documentation is written for human developers, but an increasing share of software changes is now authored by autonomous coding agents. Which documents they consult, when, and what follows remain unknown. We conduct a behaviour-grounded study of agent-documentation interaction across two public datasets: 557 agentic coding sessions from SWE-chat, yielding 94,813 development events including 3,033 documentation interactions; and 33,097 agentic pull requests from AIDev, with 690,260 classified file-level change records. Four findings challenge current documentation practice. First, agents' documentation work is dominated by agent-facing artefacts: instruction files and working notes account for 60.5% of all documentation interactions, versus 10.6% for classical technical documentation and 1.3% for API references. Second, the link between consultation and code editing is unresolved: the adjacent transition probability is 0.002 and the unadjusted three-event lift 1.05, whereas a stage-adjusted model places it above unity (OR 1.33 [1.09, 1.62]); documentation creation is elevated unadjusted (lift 1.67) but its adjusted interval includes unity. Third, no explicit documentation-based validation sequence was observed, and consultation is associated with less immediate testing (lift 0.23, cluster CI 0.08-0.45; adjusted OR 0.39 [0.25, 0.60]). Fourth, consultation is self-initiated (70.2%) far more often than failure-driven (7.5%), and documentation trails code: among multi-commit pull requests changing both, code is touched first 4.7x more often. From these traces we derive a descriptive model of agent-documentation interaction as a two-lobed cycle rather than a linear journey, and show that two widely assumed properties of "agent-friendly" documentation - actionability and verifiability - lack consistent behavioural support. We release our pipeline, coding scheme, and event-level data.

Comments14 pages, 1 figure, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑