arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于聚类的平衡采样与数据并行分配用于高性能微调

Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning

Hyunjin Kim, Youngeun Nam, Jaemin Han, Wonhyeok Choi, Jae-Gil Lee

arXiv 2609.12584首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对指令微调数据冗余与不平衡问题,提出聚类感知的平衡采样框架CluSTER,通过梯度空间聚类和DP感知分配实现高效数据缩减,最多减少69.6%训练时间且几乎无精度损失。

AI 中文摘要

大型语言模型(LLM)的指令微调数据集通常规模庞大、冗余且类别不平衡,限制了高效适配。朴素的大批量微调会重复包含过度代表的样本组,而弱覆盖代表不足但信息丰富的样本组,尤其是在多GPU数据并行(DP)环境下。我们提出CluSTER,一种面向DP指令微调中高效数据缩减的聚类感知平衡采样框架。CluSTER通过梯度空间聚类和DP感知的平衡分配来策划一个具有代表性的缩减数据集,确保在聚类和工作节点之间实现双重覆盖,同时通过加权更新保持原始数据分布。因此,CluSTER减少了冗余计算并提高了训练稳定性,且不损害模型质量。在多个指令微调数据集上,与先前的采样和数据缩减方法相比,CluSTER将训练时间最多减少69.6%,且几乎无精度损失。代码可在https://this URL获取。

英文摘要

Instruction-tuning datasets for large language models (LLMs) are often large, redundant, and imbalanced, limiting efficient adaptation. Naive large-batch fine-tuning repeatedly includes overrepresented sample groups while weakly covering underrepresented but informative ones, especially under data parallelism (DP) across multiple GPUs. We propose CluSTER, a Cluster-aware balanced Sampling framework for Training Efficient data Reduction in DP instruction tuning. CluSTER curates a representative reduced dataset through gradient-space clustering and DP-aware balanced allocation, ensuring dual-level coverage across clusters and workers, while preserving the original data distribution by weighted update. As a result, CluSTER reduces redundant computation and improves training stability without compromising model quality. Across multiple instruction-tuning datasets, CluSTER reduces training time by up to 69.6% with almost no accuracy loss compared to prior sampling and data reduction methods. Code is available at https://github.com/kaist-dmlab/CluSTER.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑