arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SDO:面向高效大语言模型后训练的结构感知数据组织

SDO: Structure-Aware Data Organization for Efficient LLM Post-Training

Jinliang Gao, Ning Yang, Hai Wang, Baili Xiao, Pin Lyu

arXiv 2607.27273首次发表:更新:

AI 中文总结

针对大语言模型后训练中数据组织固定导致的样本优化失衡问题,提出SDO框架,通过逐轮调整小批量与样本暴露,加快多任务收敛并均衡不同问题类型的准确率。

AI 中文摘要

大语言模型的后训练成本高昂,现有效率提升方法主要聚焦于选择信息丰富的样本或设计训练调度方案。然而,数据组织本身通常被视为静态预处理步骤:基于嵌入的分组方法会在训练前构建固定划分,无法适配优化过程中不断变化的样本暴露情况。因此,尽管不同样本的优化需求存在差异,但所有样本获得的暴露程度相近,导致部分样本存在冗余更新,而另一些样本则优化不足。为解决这一问题,我们提出SDO(Structure-Aware Data Organization,结构感知数据组织),这是一种即插即用的数据组织框架,具备基于暴露的反馈机制,可根据表示空间结构组织小批量数据组成与样本暴露情况。SDO基于冻结的外部嵌入逐轮运行,避免了模型预热训练的开销:在每一轮内,基于 locality(局部性)的批处理通过KNN邻域遍历形成连贯的小批量数据;在各轮之间,基于暴露平衡的调度方案记录每个样本的参与情况,并降低过度暴露样本的采样概率,以保持长期覆盖度。在SFT、DPO和GRPO三类任务中,SDO均能加快收敛速度,在训练的早中期阶段增益最为显著,可生成更连贯的梯度,使不同问题类型的准确率更均衡,且不会永久排除训练样本。

英文摘要

Post-training of large language models is expensive, and existing efficiency improvements mainly focus on selecting informative samples or designing training schedules. However, data organization itself is usually treated as a static preprocessing step: embedding-based grouping methods construct fixed partitions before training and cannot adapt to the evolving sample exposure during optimization. As a result, all samples receive similar exposure despite their different optimization needs, leading to redundant updates for some samples while leaving others under-optimized. To address this problem, we propose SDO (Structure-Aware Data Organization), a plug-and-play data organization framework with an exposure-driven feedback mechanism that organizes mini-batch composition and sample exposure according to representation-space structure. SDO operates epoch by epoch on frozen external embeddings, avoiding model warm-up training overhead: within each epoch, locality-aware batching forms coherent mini-batches via KNN neighborhood traversal; across epochs, exposure-balanced scheduling records per-sample participation and reduces the sampling probability of over-exposed samples to preserve long-term coverage. Across SFT, DPO, and GRPO, SDO accelerates convergence, with the largest gains observed in the early-to-mid phase, producing more coherent gradients and more balanced accuracy across question types without permanently excluding training samples.

Comments9 pages, 5 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑