发表机构
Sun Yat-sen University; Shenzhen Loop Area Institute; The Chinese University of Hong Kong, Shenzhen; Shenzhen Research Institute of Big Data(中山大学; 深圳河套学院; 香港中文大学(深圳); 深圳市大数据研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出解构视角,将代码语料按计算模式分类,通过微调实验发现代码数据对模型和任务的影响因类别而异,且紧凑混合优于全语料,体现“少即是多”。
AI 中文摘要
将代码数据作为一个整体语料库进行评估,可能会掩盖哪些类型的代码数据对哪些模型和下游任务有益。有效的数据选择需要理解各个类别的益处,以及这些益处是否在类别组合时仍然存在。我们引入一种解构视角来研究LLM后训练中的这些效应。我们首先根据解决方案的计算模式,将一个经过执行验证的代码语料库分解为可解释的类别。通过受控微调实验,我们在指令调优模型上,针对问答、数学和代码生成任务,比较了各个类别与平衡混合的效果。得到的响应图显示,平均问答性能有重复性的提升,而同一类别可能改善一个模型或任务,却损害另一个。表现最佳的类别也随起始模型和目标任务而变化。然后,我们根据这些结果构建紧凑的混合,并检查在联合训练下,单个类别中观察到的益处是否持续存在。在选定的模型-任务对上,其组成成分各自都能改善目标任务的混合,优于其最佳成分和全语料训练,同时仅使用全语料的约10-15%。这些探索性发现展示了一种“少即是多”的模式,并强调了后训练中代码数据的价值取决于为哪个模型和任务组合哪些类别。
英文摘要
Evaluating code data as a single corpus can obscure which types of code data benefit which models and downstream tasks. Effective data selection requires understanding both the benefits of individual categories and whether these benefits persist when categories are combined. We introduce a decompositional lens for studying these effects in LLM post-training. We first decompose an execution-verified code corpus into interpretable categories based on the computational patterns of its solutions. Through controlled fine-tuning experiments, we compare individual categories with a balanced mixture across instruction-tuned models on question answering, mathematics, and code generation. The resulting response maps reveal recurring gains in average question-answering performance, while the same category can improve one model or task and degrade another. The best-performing category also varies with the starting model and target task. We then compose compact mixtures guided by these results and examine whether benefits observed in individual categories persist under joint training. On selected model--task pairs, mixtures whose constituents each improve the target task outperform both their best constituent and full-corpus training while using roughly 10--15\% of the full corpus. These exploratory findings illustrate a \emph{less is more} pattern and highlight how the value of code data in post training depends on which categories are combined for which model and task.