arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RepoProbe:基于清单的架构感知代码仓库理解基准测试

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists

Yuexi Yang, Alyssa Wu, Ji Luo, Richeng Xuan, Zhichao Hu, Yuhong Liu, Zhen Qin

arXiv 2608.04783首次发表:更新:

发表机构

Hunyuan, Tencent; Zhejiang University; Ningbo Global Innovation Center, Zhejiang University(腾讯混元; 浙江大学; 浙江大学宁波全球创新中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出RepoProbe基准,采用基于清单的验证协议,发现SOTA LLMs存在编辑偏差,且高清晰度与技术正确性间有差距,提升了代码仓库理解评估的可靠性。

AI 中文摘要

大型语言模型(LLMs)在软件工程中的应用,已将重点从函数级生成转向代码仓库级别的辅助。然而,现有基准大多依赖GitHub Issues中的bug报告,这往往让模型通过对错误日志的模式匹配绕过真正的理解。这种不一致低估了编辑偏差,即模型过早提出代码修改而非理解现有代码仓库架构的情况。此外,当前以LLM作为评判者的标量评分存在高方差和低可解释性问题。本研究推出RepoProbe,这是一个通过GitHub Discussions开展的开放式问答,用于评估代码仓库级代码理解的新型基准,聚焦开放式架构查询而非缺陷报告。为确保严格评估,我们提出基于清单的验证协议,将答案分解为可验证的原子事实,从而用客观验证替代主观评分。我们对最先进(SOTA)LLMs的评估显示,其在高清晰度与基于证据的技术正确性之间存在持续差距,还定量证实了编辑偏差的普遍性,即模型优先进行代码生成而非架构分析。最后,我们证明与传统标量评分评估相比,我们的验证协议显著提升了评估可靠性。

英文摘要

The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repository-scale assistance. However, existing benchmarks largely rely on bug reports from GitHub Issues, which often allow models to bypass genuine understanding via pattern matching on error logs. This misalignment under-measures Edit Bias, which refers to premature generation, where models prematurely propose code modifications instead of understanding the existing repository architecture. Furthermore, current LLM-as-a-Judge scalar scoring suffers from high variance and low interpretability. This work introduces RepoProbe, a novel benchmark for evaluating repository-level code understanding through open-ended Q&A using GitHub Discussions, which focuses on open-ended architectural inquiries rather than defect reporting. To ensure rigorous evaluation, we propose a Checklist-Based Verification Protocol that decomposes answers into atomic, verifiable facts, thereby replacing subjective ratings with objective verification. Our evaluation of state-of-the-art (SOTA) LLMs reveals a persistent gap between high clarity and evidencegrounded technical correctness. It also quantitatively confirms the prevalence of edit bias, in which models prioritize code generation instead of architectural analysis. Finally, we demonstrate that our verification protocol significantly improves evaluation reliability compared to traditional evaluations with scalar scoring.

CommentsAccepted to the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026). Replication package: https://github.com/Tencent-Hunyuan/RepoProbe

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑