arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27891cs.SEcs.AI

Schrödinger的代码仓库:LLM是学会了SWE-bench还是记住了它?

Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?

Silin Chen, Yufei Yang, Xiaodong Gu, Yuling Shi, Chengcheng Wan, Haibing Guan

首次发表
浏览论文内容

中文总结 AI 辅助

针对仓库级编码基准的数据泄露问题,提出SchrodingerRepo动态实例化评估框架,通过四种转换消除熟悉线索,实验显示LLM性能下降且交互成本上升,表明智能体部分依赖记忆而非推理。

中文摘要 AI 辅助

仓库级编码基准已成为评估编码智能体的标准,然而它们固有地存在数据泄露问题,因为这些基准构建于广泛用于训练的流行开源仓库之上。因此,优异的表现可能反映的是对典型仓库线索的记忆,而非稳健的仓库推理能力。我们提出了SchrodingerRepo(薛定谔仓库),一个在动态实例化的仓库表示下测试编码智能体的评估框架。SchrodingerRepo不是反复使用测试仓库的静态表示,而是将测试仓库视为一个评估时的潜在变量,仅在智能体进入评估环境时才动态实例化。实例化的仓库保留了原始可执行行为,同时通过四个转换级别削弱熟悉的线索,如命名约定、文件布局和实现模式:问题陈述重构、命名空间重映射、文件内布局重排和保持功能的代码重写。我们在SWE-bench Verified和SWE-QA上评估了流行的LLM。结果表明,移除熟悉的仓库线索持续降低智能体性能,并显著增加各模型的交互成本。进一步分析显示,额外成本主要由仓库探索和定位难度的增加引起。这些发现表明,当前编码智能体可能部分依赖记忆的仓库侧线索,强调了在动态实例化的仓库表示下进行评估的必要性。

英文摘要

Repository-level coding benchmarks have become the standard for evaluating coding agents, yet they inherently suffer from data leakage because they are built upon popular open-source repositories repeatedly used for training. Consequently, strong performance may reflect memorization of canonical repository cues rather than robust repository reasoning. We propose SchrodingerRepo (Schrödinger's Repository), an evaluation framework for testing coding agents under dynamically instantiated repository representations. Instead of repeatedly using a static representation of the test repository, SchrodingerRepo treats the test repository as an evaluation-time latent variable that is dynamically instantiated only when the agent enters the evaluation environment. The instantiated repository preserves the original executable behavior while eroding familiar cues such as naming conventions, file layouts, and implementation patterns through four transformation levels: problem statement reconstruction, namespace remapping, intra-file layout reordering, and functionality-preserving code rewriting. We evaluate popular LLMs on SWE-bench Verified and SWE-QA. Results show that removing familiar repository cues consistently degrades agent performance and substantially increases interaction costs across models. Further analysis reveals that the additional cost is primarily caused by increased difficulty in repository exploration and localization. These findings suggest that current coding agents may partially rely on memorized repository-side cues, highlighting the need for evaluation under dynamically instantiated repository representations.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Xi’an Jiaotong University(西安交通大学)
  • East China Normal University(华东师范大学)
  • Shanghai Innovation Institute(上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑