arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25241cs.SEcs.AI

几页Markdown:采用编码智能体后的已提交AI配置与更低质量成本

A Few Pages of Markdown: Committed AI Configuration and Lower Quality Cost after Coding-Agent Adoption

Yegor Denisov-Blanch, Shyam Agarwal, Pavel Azaletskiy, Hao He, Rylan Schaeffer, Brando Miranda, Bogdan Vasilescu, Sanmi Koyejo

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出RAMP四级AI成熟度模型,在441个代码库中验证其有效性,发现编码智能体可加速开发,且提交AI配置能降低认知复杂度与静态分析警告增幅。

中文摘要 AI 辅助

编码智能体可提升开发速度,但也会增加技术债务。现有研究仅报告了采用者的平均效应,掩盖了团队间的巨大差异。我们提出RAMP(Repository AI Maturity Profile,即代码库AI成熟度档案),这是一个基于团队提交用于配置AI工具的版本控制制品的四级累积成熟度模型。RAMP的范围从行为规则和编码标准,到命名智能体定义,再到多智能体编排,观测到的实践集中在前三个级别。在441个代码库中,这些级别表现为累积量表,独立人工标注在保留样本上重现了RAMP代码库级标签的97%。采用过程是累积、单向且设置后不再变更的:73.8%的制品仅提交一次且从未修改。在每个层级内重新评估现有智能体采用面板,发现智能体无论成熟度如何都能加速开发(提交量增加28%-38%),但质量存在差异:在能识别出对比的智能体优先代码库中,未提交AI配置的代码库的认知复杂度增幅约为提交配置代码库的两倍(53%对比27%),静态分析警告的增幅为1.7倍。由于成熟度是观测性的,相关的工程规范或模型能力可能是该差距的部分原因;我们将这些发现作为假设生成,并发布RAMP作为可复用工具。

英文摘要

Coding agents increase development velocity but also technical debt. Prior work reports only average effects across adopters, hiding wide differences between teams. We introduce RAMP (Repository AI Maturity Profile), a four-level cumulative maturity model grounded in version-controlled artifacts that teams commit to configure AI tools. RAMP runs from behavioral rules and coding standards through named agent definitions to multi-agent orchestration, with observed practice concentrated in the first three levels. Across 441 repositories the levels behave as a cumulative scale, and independent human annotation reproduces RAMP's repository-level labels on 97% of a held-out sample. Adoption is cumulative, forward-only, and set-and-forget: 73.8% of artifacts are committed once and never modified. Re-estimating an existing agent-adoption panel within each stratum, agents accelerate development regardless of maturity (28-38% more commits), but quality diverges: among agent-first repositories, where the contrast is identified, those without committed AI configuration show roughly twice the increase in cognitive complexity (+53% versus +27%) and 1.7x the increase in static-analysis warnings. Because maturity is observational, correlated engineering discipline or model capability may explain part of the gap; we present these findings as hypothesis-generating and release RAMP as a reusable instrument.

发表机构

  • Stanford University(斯坦福大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Grid Dynamics

机构由 AI 辅助整理,请以论文原文为准。

↑