arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04208cs.SE

AI生成代码,人类承担债务:智能体生成代码的可持续性与演化实证研究

AI Writes Code, Humans Pay the Debt. An Empirical Study on the Sustainability and Evolution of Agent-Generated Code

发表机构南丹麦大学 · 奥卢大学 · 夏威夷大学
查看机构详情
  • University of Southern Denmark(南丹麦大学)
  • University of Oulu(奥卢大学)
  • University of Hawaii(夏威夷大学)

机构由 AI 辅助整理,请以论文原文为准。

Antonino Coppola, Matteo Esposito, Rick Kazman, Valentina Lenarduzzi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过SQuaD数据集对比智能体生成代码与开发者实际提交代码,探究智能体生成代码对软件质量、技术债务及演化的影响,为LLM智能体的负责任应用提供依据。

中文摘要 AI 辅助

背景:生成式AI编码智能体在软件工程中的应用日益广泛,正改变开发者实现与维护代码的方式。这些系统虽能带来短期生产力提升,但其对软件质量和技术债务的长期影响仍不明确。目标:本研究旨在探究智能体生成代码如何影响软件质量,重点关注问题定位准确率、技术债务引入情况及其随时间的演化。方法:我们将开展一项基于软件仓库挖掘的大规模研究,使用SQuaD数据集,涵盖62.8万张问题工单。我们将为这些工单生成基于智能体的实现方案,再通过静态分析指标与工具,将其与开发者实际提交的代码进行对比。我们将分析提交层面及各版本间的差异,并采用系统基准测试策略选取多个基于大语言模型(LLM)的智能体进行研究。预期结果:我们期望提供智能体开发模式带来的权衡的实证证据,包括定位准确率的差异、技术债务引入情况的变化,以及长期演化可能存在的分歧。我们希望研究结果能凸显不同LLM间的差异,加深对智能体参与下软件演化的理解,并为软件开发中更负责任地采用智能体提供依据。

英文摘要

Context. The increasing adoption of Generative AI coding agents in software engineering is transforming how developers implement and maintain code. While these systems provide short-term productivity benefits, their long-term impact on software quality and technical debt remains unclear. Aim. We aim to investigate how agent-generated code affects software quality, focusing on issue localization accuracy, the introduction of technical debt, and its evolution over time. Method. We will conduct a large-scale mining software repositories study using the SQuaD dataset, employing a candidate set of 628k issue tickets. We will generate agent-based implementations for these issues, and compare them with the actual commits done by developers using static analysis metrics and tools. We will analyze differences at the commit level and across releases, and we will consider multiple LLM-based Agents selected through a systematic benchmarking strategy. Expected Results. We expect to provide empirical evidence on the trade-offs introduced by agent-based development, including differences in localization accuracy, variations in technical debt introduction, and potential divergence in long-term evolution. We expect the results to highlight variability across LLMs, to enrich our understanding of software evolution with Agents, and to inform more responsible adoption of Agents in software development.

补充信息

↑