DeepRepoQA:基于深度智能体探索的代码仓库问答系统
DeepRepoQA: Code Repository Question Answering with Deep Agent Exploration
浏览论文内容
中文总结 AI 辅助
针对现有代码仓库问答方法缺乏深度推理能力的问题,提出基于LLM智能体与蒙特卡洛树搜索的DeepRepoQA框架,在SWE-QA基准上实现了性能显著提升。
中文摘要 AI 辅助
回答开发者关于软件仓库的问题是软件工程领域一个关键但未被充分探索的问题。尽管现有的仓库理解方法推动了该领域的发展,但它们主要依赖表层代码检索,缺乏对多文件、复杂软件架构进行深度推理,以及基于长距离代码依赖关系确定答案的能力。为解决这些局限,我们提出DeepRepoQA,一个用于仓库级代码理解的新型问答框架。DeepRepoQA基于智能体框架构建,其中大语言模型(LLM)智能体通过对仓库结构进行系统的树搜索来寻找答案。采用蒙特卡洛树搜索(MCTS)机制,使智能体能够动态搜索、导航和检查代码,从而实现对长距离代码依赖关系的有效多跳推理。在SWE-QA基准上开展的综合实验表明,DeepRepoQA相比强基线方法取得了显著的性能提升,验证了MCTS引导的系统探索对于多跳仓库推理的有效性。
英文摘要
Answering developer questions about a software repository is a critical yet under-explored problem in software engineering. While existing repository understanding methods have advanced the field, they predominantly rely on surface-level code retrieval and lack the ability for deep reasoning over multiple files, complex software architectures, and grounding answers in long-range code dependencies. To address these limitations, we propose DeepRepoQA, a novel question answering (QA) framework for repository-level code understanding. DeepRepoQA builds on an agentic framework where LLM agents find answers through a systematic tree search over the repository structure. A Monte-Carlo Tree Search (MCTS) mechanism is employed to empower agents to dynamically search, navigate, and inspect code, enabling effective multi-hop reasoning over long-range code dependencies. Comprehensive experiments on the SWE-QA benchmark demonstrate substantial performance gains over strong baselines, validating the effectiveness of systematic MCTS-guided exploration for multi-hop repository reasoning.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- The Hong Kong University of Science and Technology(香港科技大学)
- University of California San Diego(加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。