arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAP4CS:用于高效代码检索器微调的多维数据剪枝框架

MAP4CS: A Multi-dimensional Data Pruning Framework for Efficient Code Retriever Fine-tuning

Yuxuan Chen, Mingwei Liu, Guangsheng Ou, Zekai Zhang, Zike Li, Yanlin Wang, Pelin Zheng

arXiv 2610.11727首次发表:更新:

发表机构

School of Software Engineering, Sun Yat-sen University(中山大学软件学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MAP4CS框架,通过多维数据剪枝仅用5%训练数据实现高效代码检索器微调,性能优于随机采样,可媲美甚至优于全量数据微调,验证了数据中心AI的“少即是多”假设。

AI 中文摘要

检索增强生成(RAG)已成为软件工程领域增强大型语言模型(LLMs)领域特定知识的核心技术。然而,由于大规模代码语料库固有的噪声和冗余,将检索器适配到不断演进的代码仓库仍具挑战性。对完整语料库进行标准微调计算成本高昂,且常因低质量样本的负迁移导致性能欠佳;而简单的随机采样无法保证数据的代表性。为应对这些挑战,本文提出MAP4CS(面向代码搜索的多维感知剪枝),这是一种自适应数据剪枝框架。MAP4CS通过整合句法结构、语义多样性和分布表示,结合严格的基于规则的过滤流程,识别出小型高质量核心子集。在两个大规模数据集上的大量实验表明,MAP4CS仅使用5%的训练数据,始终优于随机采样基线;其性能可与完整数据集微调相媲美,甚至更优,验证了数据中心AI中的“少即是多”假设。此外,语言分析揭示了一种自适应优化机制:MAP4CS自动对冗余语料库执行去重,对混乱语料库执行去噪,构建出词汇多样且信息密集的训练语料库。

英文摘要

Retrieval-Augmented Generation (RAG) has become a cornerstone in software engineering for enhancing Large Language Models (LLMs) with domain-specific knowledge. However, adapting retrievers to evolving code repositories remains challenging due to the noise and redundancy inherent in massive code corpora. Standard fine-tuning on the full corpus is computationally expensive and often leads to sub-optimal performance due to negative transfer from low-quality samples. Conversely, simple random sampling fails to guarantee data representativeness. To address these challenges, we propose MAP4CS (Multi-dimensional Awareness Pruning for Code Search), an adaptive data pruning framework. MAP4CS identifies a small, high-quality core subset by integrating syntactic structure, semantic diversity, and distributional representation, followed by a rigorous rule-based filtering pipeline. Extensive experiments on two large-scale datasets demonstrate that MAP4CS consistently outperforms random sampling baselines using only 5% of the training data. Remarkably, it achieves performance comparable to, or even superior to, fine-tuning on the full dataset, validating the ''less is more'' hypothesis in data-centric AI. Furthermore, linguistic analysis reveals an adaptive optimization mechanism: MAP4CS automatically functions as a de-duplicator for redundant corpora and a denoiser for chaotic ones, constructing a training corpus that is both lexically diverse and information-dense.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑