arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MRCoder:一种面向仓库级代码生成的高效上下文选择方法

MRCoder: An Efficient Context Selecting Approach for Repository-Level Code Generation

Peiding Wang, Li Zhang, Fang Liu

arXiv 2607.26805首次发表:更新:

AI 中文总结

MRCoder是一种基于Map-Reduce范式的仓库级代码生成上下文选择框架,通过SADGS策略和并行验证,在提升代码生成准确率的同时降低了token消耗与推理时间。

AI 中文摘要

大型语言模型(LLM)已展现出强大的代码生成能力,但仓库级代码生成仍具挑战性,因为它需要有效识别和利用仓库特定上下文。检索增强生成(RAG)虽会纳入相关代码片段,却常引入冗余上下文干扰LLM利用有效信息,导致生成质量下降、计算成本增加。此外,现有上下文选择与压缩方法难以平衡效率与质量,要么引入额外计算开销,要么无法有效选择有效上下文。本文提出MRCoder,一种可提升仓库级代码生成有效性与效率的高效上下文选择框架。MRCoder采用Map-Reduce范式:在Map阶段,轻量级草稿模型对分区上下文生成草稿,结构感知草稿引导选择(SADGS)通过API一致性与逻辑相似性,基于草稿选择信息丰富的上下文;在Reduce阶段,将优化后的上下文聚合用于最终生成,并行验证策略进一步加速解码。我们在CoderEval和DevEval两个广泛使用的仓库级代码生成基准上,以Qwen2.5-Coder和DeepSeek-Coder作为骨干LLM评估MRCoder。实验结果表明,MRCoder相比强基线提升了代码生成准确率,同时将token消耗降低30%至50%,推理时间最多缩短52%。这些结果证明,本文提出的结构化草稿引导上下文选择策略,对提升仓库级代码生成的质量与效率至关重要。

英文摘要

Large language models (LLMs) have demonstrated strong capabilities in code generation. However, repository-level code generation remains challenging, as it requires effectively identifying and utilizing repository-specific context. While retrieval-augmented generation (RAG) incorporates relevant code snippets, it often introduces redundant context that interferes with the LLM's ability to utilize relevant information, leading to degraded generation quality and increased computational cost. Moreover, existing context selection and compression methods struggle to balance efficiency and quality, either introducing additional computational overhead or failing to effectively select valid context. In this paper, we propose MRCoder, an efficient context selection framework that improves both the effectiveness and efficiency of repository-level code generation. MRCoder adopts a Map-Reduce paradigm: in the Map Phase, a lightweight draft model generates drafts over partitioned contexts, and Structure-Aware Draft-Guided Selection (SADGS) selects informative contexts based on drafts through API consistency and logical similarity; in the Reduce Phase, the refined contexts are aggregated for final generation, with a parallel verification strategy further accelerating decoding. We evaluate MRCoder on two widely used repository-level code generation benchmarks, CoderEval and DevEval, using Qwen2.5-Coder and DeepSeek-Coder as backbone LLMs. Experimental results show that MRCoder improves code generation accuracy over strong baselines while reducing token consumption by 30 to 50% and inference time by up to 52%. These results demonstrate that our proposed structured and draft-guided context selection strategy is crucial for improving both the quality and efficiency of repository-level code generation

CommentsUnder review in TOSEM

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑