arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

训练感知的目标覆盖用于合成数据选择

Training-Aware Target Coverage for Synthetic Data Selection

Yang Ba, Michelle V. Mancenido, Rong Pan

arXiv 2610.00814首次发表:更新:

发表机构

Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出训练感知的目标覆盖(TATC)方法,基于线性理论量化合成数据对目标任务的边际价值,在微调中通过选择有益候选并扩展目标相关覆盖,在数学推理任务上优于其他选择方法。

AI 中文摘要

合成数据越来越多地用于扩展大型语言模型的训练,但更多的合成数据并不一定产生更好的模型。有用的合成数据必须添加与目标任务相关的信息,而不引入抵消其益处的错误,并且一个示例的价值会随着训练集的增长而变化。我们开发了一个线性理论来刻画这种权衡,并确定合成数据在何处有用、应添加多少,以及向现有集合添加一个示例的边际价值。分析显示了仅输入覆盖就足够的情况,以及必须同时考虑合成错误的情况。在这些结果的指导下,我们引入了训练感知的目标覆盖(TATC),一种用于大型语言模型微调的合成数据选择方法。TATC识别出训练效果对目标任务有益的候选样本,并在其中进行选择,以扩展目标相关方向(现有数据尚未覆盖的方向)的覆盖。在文本和图像数据上的实验验证了该线性理论。在数学推理任务中,TATC选择合成解来微调Qwen2.5-Math-1.5B-Instruct,并在GSM8K上跨选择预算优于其他合成数据选择方法。总之,我们通过量化和最大化合成数据对目标任务的价值,提供了一种原则性的合成数据选择方法。

英文摘要

Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the training set grows. We develop a linear theory that characterizes this tradeoff and determines where synthetic data are useful, how much should be added, and the marginal value of adding one example to an existing set. The analysis shows the conditions when input coverage alone is sufficient and when synthetic errors must also be considered. Guided by these results, we introduce \emph{Training-Aware Target Coverage} (TATC), a synthetic data selection method for LLM fine-tuning. TATC identifies candidates whose training effects are beneficial to the target task and selects among them to expand coverage of target-relevant directions not already represented by the available data. Experiments on text and image data verify the linear theory. With mathematical reasoning tasks, TATC selects synthetic solutions for fine-tuning Qwen2.5-Math-1.5B-Instruct and outperforms alternative synthetic-data selection methods on GSM8K across selection budgets. In summary, we provide a principled approach to synthetic data selection by quantifying and maximizing its value to the target task.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑