发表机构
LightOn; Inria Paris; University of Copenhagen(LightOn; 法国国家信息与自动化研究所巴黎分部; 哥本哈根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出使用ColGREP语义搜索工具训练小型专用智能体,以高效定位代码文件,在准确性、延迟和令牌效率上优于传统GREP方法,并支持跨语言泛化。
AI 中文摘要
从自然语言请求中定位相关文件是代码库上运行的智能体的核心子任务。我们研究是否可以将此任务委托给紧凑的专用模型,以实现设备端搜索,同时减少较大智能体的令牌使用、延迟和推理成本。我们表明,语义搜索改善了文件定位,在准确性、跨语言迁移和推理效率方面都有所提升。为了研究这一设置,我们引入了一个针对文件定位智能体的训练框架,该框架围绕ColGREP构建,ColGREP是一种基于后期交互检索模型的本地语义搜索工具。我们的方法结合了在教师轨迹上的加权监督微调(根据检索结果分配回合级信用)与基于定位质量的强化学习。我们训练了三个参数少于20亿的模型族,用于制定搜索查询、检查检索到的内容并识别相关文件。在从SWE-bench Lite和Multi-SWE-bench Flash派生的定位任务上,配备ColGREP的智能体显著优于其基础模型,并超过了相应的基于GREP的智能体。除了提高定位准确性外,ColGREP在CPU上将平均端到端轨迹延迟降低了44.1%,同时使用的令牌减少了29.1%,并且能够更好地泛化到微调期间未见过的编程语言。这些结果表明,紧凑的、工具专用的定位智能体可以为自然语言请求与大型代码库之间提供高效的接口。
英文摘要
Locating relevant files from natural-language requests is a core subtask for agents operating over code repositories. We investigate whether this task can be delegated to compact, specialized models to enable on-device search while reducing the token usage, latency, and inference cost of larger agents. We show that semantic search improves file localization, with gains in accuracy, cross-language transfer, and inference efficiency. To study this setting, we introduce a training framework for file-localization agents built around ColGREP, a local semantic search tool based on late-interaction retrieval models. Our recipe combines weighted supervised fine-tuning on teacher trajectories, assigning turn-level credit based on retrieval outcomes, with reinforcement learning on localization quality. We train three model families with fewer than two billion parameters to formulate search queries, inspect retrieved content, and identify relevant files. On localization tasks derived from SWE-bench Lite and Multi-SWE-bench Flash, ColGREP-equipped agents substantially improve over their base models and outperform corresponding GREP-based agents. In addition to improving localization accuracy, ColGREP reduces mean end-to-end trajectory latency by 44.1\% on CPU while using 29.1\% fewer tokens, and enables better generalization to programming languages unseen during fine-tuning. These results suggest that compact, tool-specialized localization agents can provide an efficient interface between natural-language requests and large codebases.