发表机构
College of Computing and Data Science; Nanyang Technological University(计算与数据科学学院; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对数据选择中元学习方法的细粒度评估与可迁移性权衡问题,提出基于逐点值匹配目标的可扩展框架TESS,在LLM安全性和指令调优上实现跨数据集与跨模型规模的有效迁移。
AI 中文摘要
数据选择对于在庞大且异构的语料库上训练大型语言模型至关重要。用于训练数据选择的元学习提供了一种有原则的替代启发式评分的方法,它通过从目标验证目标中学习数据权重来实现,但现有方法在细粒度评估和对未见数据的可迁移性之间面临权衡。一个自然的解决方案是用选择网络替代每样本权重。然而,我们发现将此类网络直接纳入现有的元学习训练数据选择目标会导致优化不稳定和泛化性能差,其原因在于权重抑制和对易学习特征的持续依赖。为了解决这些问题,我们提出了可迁移示例评分与选择(TESS),这是一个基于逐点值匹配目标(PVM)的可扩展数据选择框架。在大型语言模型安全性和定向指令调优上的实验表明,该方法在数据集之间(从子集到完整语料库,以及从较小模型到较大模型)具有强大的迁移能力。
英文摘要
Data selection is critical for training large language models on massive and heterogeneous corpora. Meta-learning for Training-data Selection offers a principled alternative to heuristic scoring by learning data weights from a target validation objective, but existing methods face a trade-off between fine-grained valuation and transferability to unseen data. A natural solution is to replace per-sample weights with a selection network. However, we find that directly incorporating such a network into existing MTS objectives leads to unstable optimization and poor generalization, caused by weight suppression and persistent reliance on easy-to-learn features. To address these issues, we propose Transferable Example Scoring and Selection (TESS), a scalable data-selection framework built on a Pointwise Value Matching objective (PVM). Experiments on LLM safety and targeted instruction tuning demonstrate strong transfer across datasets, from subsets to full corpora, and from smaller to larger models.