AI 中文总结
研究文件级漏洞定位中代码表示的作用,在LCA和SWE数据集上比较五种代码表示,通过实验发现角色感知摘要性价比最佳,结合互补结果和LLM排序可进一步提升效果,案例研究验证技术有效性,强调代码表示应是智能体定位管道的重要设计选择。
AI 中文摘要
基于大语言模型(LLM)的智能体越来越多地用于支持软件开发,但其在仓库级任务中的性能取决于检索正确的代码上下文。现有研究探索了使用传统信息检索进行文件级定位。然而,文本代码表示在检索和定位中的作用仍未得到充分探索。我们将文件级漏洞定位研究为一个表示驱动的检索问题。在Long Code Arena(LCA)和SWE-bench Verified(SWE)数据集上,比较了五种代码表示。实验包括词汇、语义和基于LLM的检索,以及基于LLM的检索后排序。通过表示占用空间量化表示成本。发现代码表示的选择会影响定位有效性和成本。角色感知摘要在Hit@5上比文件路径表示高出40%,同时占用空间比原始源代码小10.4到20.9倍。结合互补表示结果并用LLM对检索到的候选进行排序分别可进一步提高31.9%和42.0%。总体而言,角色感知摘要提供了最佳的性价比权衡,原始源代码在某些设置下有效但成本高得多。一个无智能体的案例研究揭示了我们技术在一个知名管道中的效用,在文件定位上达到94%的Hit@6(比基线提高4.7%)。我们的发现表明,在智能体定位管道中,代码表示应作为一级设计选择,由管道阶段和成本-准确性要求指导。
英文摘要
LLM-based agents are increasingly being used to support software development, yet their performance in repository-level tasks depends on retrieving the right code context. Existing studies have explored file-level localization using traditional information retrieval over file paths and raw source code. However, the role of textual code representations in retrieval and localization remains underexplored. We study file-level bug localization as a representation-driven retrieval problem. Across the Long Code Arena (LCA) and SWE-bench Verified (SWE) datasets, we compare five code representations: file paths, raw source code, and three LLM-generated textual representations. Our experiments include lexical, semantic, and LLM-based retrieval, followed by LLM-based post-retrieval ranking. We quantify the cost incurred by a representation through the representation footprint. We find that the choice of code representation affects both localization effectiveness and cost. Role-aware summaries outperform file-path representations by up to 40% Hit@5 while requiring a representation footprint 10.4 to 20.9x smaller than raw source code. Combining complementary representation results and ranking retrieved candidates with an LLM provides further gains of up to 31.9% and 42.0%, respectively. Overall, role-aware summaries provide the best cost-effectiveness trade-off, while raw source code offers effectiveness in some settings at a significantly higher cost. A case study with Agentless reveals the utility of our techniques within a well-known pipeline, reaching 94% Hit@6 on file localization (+4.7% against the baseline). Our findings suggest that code representation should be treated as a first-class design choice in agentic localization pipelines, guided by pipeline stage and cost-accuracy requirements.