arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06640cs.SEcs.AI

刻画生产环境中AI生成C++代码的质量特征

Characterizing the Quality Profile of AI-Generated C++ in Production

发表机构谷歌
查看机构详情
  • Google(谷歌)

机构由 AI 辅助整理,请以论文原文为准。

Michael Tran, Fred Lewis, Kun Yang, Saksham Thakur, Aditya Kini, Aditya Patil, Milad Hashemi, Parthasarathy Ranganathan

首次发表
浏览论文内容

中文总结 AI 辅助

通过对大型企业2025-2026年352万次代码变更的分析,发现AI生成C++代码质量特征与人类代码不同,针对性反馈可降低其静态分析警告并提升效率。

中文摘要 AI 辅助

AI代码助手的广泛应用无疑提升了工程开发速度,但近期研究指出存在日益突出的权衡问题,揭示出代码质量与可维护性方面的持续挑战,包括前沿AI实验室在内的行业龙头企业也对此表示担忧。随着大型语言模型越来越多地被用于生成生产代码,理解其对交付软件质量的影响已成为关键优先事项,但由于可观测性障碍,在工业工作流中评估这些影响仍存在困难。本研究在一家运营着全球产品、每日服务数十亿用户的大型企业内部,研究AI生成代码对生产质量的影响。受该企业的规模和用户信任驱动,其重视代码质量,并为部署到生产环境的每一行代码构建了完善的可观测性,使我们能够克服测量障碍来评估这些影响。本研究对该企业2025年4月至2026年4月期间的AI生成C++代码进行了大规模实证分析,跟踪了其遗留代码库中的352万次代码变更,核心目的是在生产环境的大规模场景下,理解AI生成代码与人类编写代码在质量、性能和维护特征方面的差异。我们发现,AI生成的C++代码具有独特的质量特征,表现出更高的接口与耦合负担、复制与分配开销,且更依赖显式循环而非优化的标准API;这些问题转化为实际的下游成本,包括增加的评审工作量和5%-8%的计算资源消耗。不过,我们证明为模型提供针对性的、基于分类的反馈可缓解这些影响,使针对性静态分析警告减少11.1%,并提升计算效率。

英文摘要

The widespread integration of AI coding assistants offers undeniable boosts to engineering velocity. Yet, recent studies point to a growing trade-off, revealing persistent challenges with code quality and maintainability. Industry leaders, including frontier AI labs, echo these concerns. As large language models are increasingly relied upon to author production code, understanding their impact on shipped software quality has become a critical priority. However, assessing these effects in industrial workflows remains difficult due to observability barriers. We study the impact of AI-generated code on production quality within a large enterprise operating global products relied upon by billions of users daily. Driven by this scale and user trust, the organization values code quality and has built thorough observability for every line of code deployed into production, enabling us to overcome measurement barriers to assess these effects. This study presents a large-scale empirical analysis of AI-generated C++ code from April 2025 to April 2026, tracking 3.52 million code changes across this enterprise's brownfield codebase. The core purpose is to understand the quality, performance, and maintenance characteristics of AI-generated code compared to human-written code in a production environment at scale. We find that AI-generated C++ code has a distinct quality profile, showing higher rates of interface and coupling burdens, copy and allocation overheads, and a reliance on explicit loops over optimized standard APIs. These issues translate into tangible downstream costs, including increased review effort and a 5-8% increase in compute resource consumption. However, we demonstrate that providing models with targeted, taxonomy-informed feedback can mitigate these effects, leading to an 11.1% reduction in targeted static analysis warnings and improved computational efficiency.

补充信息

↑