发表机构
Peking University; Fudan University; Zhongguancun Academy; Shanghai Jiao Tong University; Tsinghua University; The Chinese University of Hong Kong, Shenzhen(北京大学; 复旦大学; 中关村学院; 上海交通大学; 清华大学; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对编码智能体在真实代码库中是否复用现有代码的问题,提出RepoReuse多轮基准,通过审计发现智能体存在复用不足和冗余累积,而通过率无法反映这些缺陷。
AI 中文摘要
编码智能体越来越多地被部署在真实代码库上进行迭代开发,然而现有的评估几乎没有回答一个基本问题:编码智能体会复用现有代码还是重新发明轮子?这个问题很重要:每一次重复实现都意味着一个修复被应用了两次,而且智能体生成代码的速度远快于人类审计的速度,因此冗余会在无人监督的情况下不断累积。为此,我们提出了RepoReuse,一个用于审计真实代码库中代码复用的多轮基准,其中需求逐轮揭示,工作区在轮次间累积。它由一个全自动流水线构建,该流水线结合了基于AST的依赖图、引导式证据收集和执行验证的任务合成,并且可以轻松扩展到新的代码库。除了通过率之外,我们还测量了复用率以及召回率和跨轮次的结构性冗余。对3000轮次的审计显示,智能体逐渐停止探索相关的代码库代码,即使其自身历史完全在工作区中,也很少复用,并且在第5轮时50.8%的任务链中留下了重复逻辑——而所有这些过程中通过率几乎不变。这些缺陷在通过率中是看不见的,这凸显了在功能正确性之外评估代码生成的必要性。
英文摘要
Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents produce code far faster than humans can audit, so redundancy accumulates unsupervised. Thus, we present \textbf{RepoReuse}, a multi-turn benchmark for auditing code reuse in real repositories, where requirements are revealed turn by turn and the workspace accumulates across turns. It is built by a fully automated pipeline combining AST-based dependency graphs, guided evidence collection, and execution-verified task synthesis, and scales readily to new repositories. Beyond pass rates, we measure the reuse rate together with recall and cross-turn structural redundancy. An audit over 3{,}000 turns shows that agents progressively stop exploring relevant repository code, reuse their own history less even when it is fully in the workspace, and leave duplicated logic in 50.8\% of task chains by turn~5---all while pass rates barely move. Such deficiencies are invisible to pass rates, underscoring the need to evaluate code generation beyond functional correctness.