RepoReasoner:评估长上下文语言模型的仓库级代码推理能力
RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models
浏览论文内容
中文总结 AI 辅助
研究针对现有代码推理基准测试局限,引入RepoReasoner评估仓库级代码推理,通过多阶段管道构建基准,评估输出预测和调用链预测能力,发现当前LLMs在仓库级推理存在局限,为未来相关研究提供方向。
中文摘要 AI 辅助
近期大型语言模型(LLMs)在软件工程任务中表现出色,但现有多数基准测试在函数级别评估代码推理,无法反映跨文件和复杂依赖结构的实际开发情况。为此引入RepoReasoner基准测试,评估仓库级代码推理的两个互补能力:输出预测和调用链预测。该基准通过多阶段管道构建,利用pytest执行的动态跟踪获取真实调用链,并基于LLM进行I/O重写以减少记忆效应。评估了七个先进LLMs,结果显示跨文件推理仍是挑战,模型在调用链预测中精度高但召回率低,且对重写数据性能下降,长上下文也未持续改善结果。这些发现凸显了当前LLMs在仓库级推理的基本局限,推动未来结构化架构理解和跨文件推理的研究。
英文摘要
Recent large language models (LLMs) have shown strong performance on software engineering tasks, yet most existing benchmarks evaluate code reasoning at the function level, where all relevant information is localized. This setting fails to reflect real-world development, which requires reasoning across multiple files and complex dependency structures. We introduce RepoReasoner, a benchmark for evaluating repository-level code reasoning. It assesses two complementary abilities: Output Prediction, which measures fine-grained, stateful execution reasoning across files, and Call Chain Prediction, which evaluates high-level architectural dependency understanding under noisy context. Our benchmark is constructed through a multi-stage pipeline that leverages dynamic tracing of pytest executions to obtain ground-truth call chains, along with LLM-based I/O rewriting to reduce memorization effects. We evaluate seven state-of-the-art LLMs. Even under oracle context, the best-performing model achieves only 69.1% Pass@1 on Output Prediction, indicating that cross-file reasoning remains a major challenge. In Call Chain Prediction, models exhibit high precision but low recall, suggesting limited multi-hop dependency understanding. Furthermore, performance drops on rewritten data reveal partial reliance on memorization, and longer contexts do not consistently improve results due to noise. These findings highlight fundamental limitations in current LLMs' repository-level reasoning and motivate future work on structured architectural understanding and cross-file inference.