CodeGrep:用于LLM编码智能体的基于强化学习训练的检索智能体
CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents
查看机构详情
- Netease Guangzhou AI Lab(网易广州人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对LLM编码智能体查找文件效率低的问题,提出基于GRPO训练的14B检索智能体CodeGrep,在SWE-Bench Verified上保持解决率的同时,减少了交互轮数和token使用量。
中文摘要 AI 辅助
现代LLM编码智能体(如Claude Code和OpenHands)存在一个共同的效率问题:它们将大量token预算用于查找要修补的文件,而非修补文件。在SWE-Bench Verified上,一个30B规模的OpenHands智能体平均每个已解决问题需要23轮交互和631K token,其中大量调用用于仓库探索期间的grep、glob和view_file操作。我们推出CodeGrep,这是一个14B规模的检索智能体,采用GRPO进行端到端训练,可发出多轮并行的grep、glob和read工具调用,并将候选文件返回给下游的冻结编码智能体。在全部500个SWE-Bench Verified实例上,CodeGrep在保持解决率的同时大幅提升了效率:与无检索基线的25.8%相比,CodeGrep的解决率为27.0%;在已解决的实例上,轮数减少15%,token减少19%。在所有检索器中,下游效用遵循一个精度阈值:精度为0.375的BM25会降低智能体性能,精度为0.445的Jina无显著影响,而精度为0.677的CodeGrep则达到了检索开始降低部署成本的阈值。为开展本研究,我们使用CATM从67K开源智能体轨迹中挖掘监督信号,并构建了用于多轮智能体强化学习的Git-worktree环境。在我们的设置中,在优势层而非奖励层应用效率信号可减少KL漂移,并清晰地转化为下游效率。我们将发布该模型、训练流水线、强化学习环境和评估工具链。
英文摘要
Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent averages 23 rounds and 631K tokens per resolved issue, with many calls spent on grep, glob, and view_file during repository exploration. We introduce CodeGrep, a 14B retrieval agent trained end-to-end with GRPO to issue multi-turn parallel grep, glob, and read tool calls and return candidate files to a frozen downstream coding agent. On all 500 SWE-Bench Verified instances, CodeGrep preserves resolve rate while substantially improving efficiency: 27.0% versus 25.8% for the no-retrieval baseline, with 15% fewer rounds and 19% fewer tokens on resolved instances. Across retrievers, downstream utility follows a precision threshold: BM25 with precision 0.375 degrades the agent, Jina with precision 0.445 is neutral, and CodeGrep with precision 0.677 crosses the threshold at which retrieval begins to reduce rollout cost. To enable this study, we mine supervision from 67K open-source agent trajectories using CATM and build a Git-worktree environment for multi-turn agent RL. In our setting, applying the efficiency signal at the advantage layer rather than the reward layer reduces KL drift and translates cleanly into downstream efficiency. We will release the model, training pipeline, RL environment, and evaluation harnesses.