代码仓库中问题定位的语义导航
Semantic Navigation for Issue Localization in Code Repository
浏览论文内容
中文总结 AI 辅助
提出SemNav框架,通过语义导航图、语义卡片和候选工作区,结合确定性检索与LLM迭代精炼,显著提升代码仓库问题定位性能。
中文摘要 AI 辅助
仓库级问题定位旨在识别并排序与解决所报告问题相关的文件和函数。LLM智能体以迭代方式处理此任务:它们识别一组潜在相关的位置,检查相应的代码,并在获取新证据时修订对这些候选的判断。然而,现有环境对此循环的支持有限:智能体必须搜索未解析的关系目标,从原始源代码重建实体语义,并在没有证据基础的情况下修订候选。为解决这些局限,我们提出SemNav,一个利用确定性检索来播种广泛候选集,并由LLM智能体持续精炼该集合的框架,从而将初始覆盖与证据引导的修订相结合。SemNav通过三个关键组件支持此过程。语义导航图通过语言服务器按需解析程序关系,实现跨文件直接导航至相关实体。问题条件语义卡片提供每个实体角色及其与问题相关性的简洁、基于源代码的解释。持久候选工作区记录每个候选及其证据基础,支持有依据的验证、修订和排序。在SWE-bench Lite和PLocBench上,SemNav优于现有基线,将File Hit@10从68.33%提升至82.67%(使用Gemma 4B)。组件消融和轨迹分析支持所有三个组件的互补作用,而语义卡片相对于全源阅读将工作上下文负载减少48.2%。SemNav还在SWE-Explore上所有七项证据质量指标中排名第一,并将下游问题解决率从44.00%提升至52.33%。
英文摘要
Repository-level issue localization aims to identify and rank the files and functions relevant to resolving a reported issue. LLM agents approach this task iteratively: they identify a set of potentially relevant locations, inspect the corresponding code, and revise their judgments about these candidates as new evidence is acquired. Existing environments, however, provide limited support for this loop: agents must search for unresolved relation targets, reconstruct entity semantics from raw source code, and revise candidates without evidential basis. To address these limitations, we present SemNav, a framework that leverages deterministic retrieval to seed a broad candidate set and an LLM agent to continually refine that set, thereby combining initial coverage with evidence-guided revision. SemNav supports this process through three key components. A Semantic Navigation Graph resolves program relations on demand through a language server, enabling direct navigation to related entities across files. Issue-conditioned Semantic Cards provide compact, source-grounded interpretations of each entity's role and relevance to the issue. A persistent Candidate Workspace records each candidate together with its evidential basis, enabling grounded verification, revision, and ranking. Across SWE-bench Lite and PLocBench, SemNav outperforms existing baselines, improving File Hit@10 from 68.33\% to 82.67\% with Gemma 4B. Component ablations and trajectory analysis support the complementary roles of all three components, while Semantic Cards reduce working-context load by 48.2\% relative to full-source reading. SemNav further ranks first on all seven evidence-quality metrics on SWE-Explore and improves downstream issue resolution from 44.00\% to 52.33\%.
发表机构
- University of Virginia(弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。