发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人工智能生成代码的开发者编辑行为,引入DECODE数据集。通过该数据集进行数据分析及对大语言模型预测代码编辑能力的基准测试,发现多数编辑在前15分钟,且微调后开源模型表现更佳,强调以开发者为中心方法对未来AI编程助手的必要性。
AI 中文摘要
人工智能生成代码中的缺陷需要软件开发人员手动修改或重新提示人工智能编程助手。手动代码编辑能提供比仅包含最终成功代码片段的Git提交更真实、更细致的编辑行为信息。由于缺乏高质量、真实的代码编辑数据,大语言模型大多基于公开的Git数据(如提交)进行训练。为填补这一空白,我们引入了DECODE(代码数据集的开发者编辑),这是一个包含53.6K条来自1000多名开发者的Python、TypeScript和JavaScript中人工智能生成代码的真实IDE代码编辑数据集。首先,我们展示了DECODE在数据分析中的效用,了解人工智能生成代码何时、为何以及如何被编辑。我们发现大多数编辑发生在接受人工智能完成后的前15分钟内,在31%的编辑轨迹中导致人工智能完成内容被删除。其次,我们用DECODE对大语言模型预测代码编辑的能力进行基准测试。我们发现基于DECODE进行微调能使开源3B模型在代码编辑预测任务上的表现显著优于前沿大语言模型。最后我们讨论了这项工作的意义,强调了以开发者为中心的机器学习方法对未来人工智能编程助手的必要性。
英文摘要
Imperfections in AI-generated code require that software developers modify the generated code manually, or by re-prompting an AI programming assistant. Manual code edits provide more realistic and granular information on editing behavior than Git commits, which only contain final successful code snippets. Yet, due to a lack of high-quality, realistic code editing data, LLMs are mostly trained on publicly available Git data (e.g., commits). To address this gap, we introduce DECODE (Developer Edits of Code Dataset), a dataset of 53.6K real-world in-IDE code edits of AI-generated code in Python, TypeScript, and JavaScript, sourced from 1K+ developers. First, we demonstrate the utility of DECODE for data analysis, obtaining insights on when, why, and how AI-generated code is edited. We find that most edits occur within the first 15 minutes after accepting an AI completion, resulting in the removal of AI completions in 31% of edit trajectories. Second, we use DECODE to benchmark the ability of LLMs to predict code edits. We find that finetuning on DECODE enables open-source 3B models to perform code edit prediction tasks significantly better than frontier LLMs. We then discuss implications of this work, emphasizing the necessity of developer-centric machine learning approaches for future AI programming assistants.