发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对指令微调数据冗余与不平衡问题,提出聚类感知的平衡采样框架CluSTER,通过梯度空间聚类和DP感知分配实现高效数据缩减,最多减少69.6%训练时间且几乎无精度损失。
AI 中文摘要
大型语言模型(LLM)的指令微调数据集通常规模庞大、冗余且类别不平衡,限制了高效适配。朴素的大批量微调会重复包含过度代表的样本组,而弱覆盖代表不足但信息丰富的样本组,尤其是在多GPU数据并行(DP)环境下。我们提出CluSTER,一种面向DP指令微调中高效数据缩减的聚类感知平衡采样框架。CluSTER通过梯度空间聚类和DP感知的平衡分配来策划一个具有代表性的缩减数据集,确保在聚类和工作节点之间实现双重覆盖,同时通过加权更新保持原始数据分布。因此,CluSTER减少了冗余计算并提高了训练稳定性,且不损害模型质量。在多个指令微调数据集上,与先前的采样和数据缩减方法相比,CluSTER将训练时间最多减少69.6%,且几乎无精度损失。代码可在https://this URL获取。
英文摘要
Instruction-tuning datasets for large language models (LLMs) are often large, redundant, and imbalanced, limiting efficient adaptation. Naive large-batch fine-tuning repeatedly includes overrepresented sample groups while weakly covering underrepresented but informative ones, especially under data parallelism (DP) across multiple GPUs. We propose CluSTER, a Cluster-aware balanced Sampling framework for Training Efficient data Reduction in DP instruction tuning. CluSTER curates a representative reduced dataset through gradient-space clustering and DP-aware balanced allocation, ensuring dual-level coverage across clusters and workers, while preserving the original data distribution by weighted update. As a result, CluSTER reduces redundant computation and improves training stability without compromising model quality. Across multiple instruction-tuning datasets, CluSTER reduces training time by up to 69.6% with almost no accuracy loss compared to prior sampling and data reduction methods. Code is available at https://github.com/kaist-dmlab/CluSTER.