arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型后训练中基于流形覆盖与稀疏特征覆盖的分层数据选择

Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training

Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu

arXiv 2608.16927首次发表:更新:

AI 中文总结

本文提出MASS算法,将LLM后训练的数据选择建模为分层覆盖问题,在Vision Flan与LLaVA-CoT上仅用少量数据即可达到或超过全量训练效果,且优于现有基线方法。

AI 中文摘要

随着监督微调数据规模不断扩大,从庞大候选池中选择高价值子集对降低训练成本、提升模型性能至关重要。现有方法常直接在原始嵌入空间中衡量多样性,该空间的几何指标会混淆主导语义方向、细粒度监督差异与局部噪声。本文将数据选择问题建模为由粗到细的分层覆盖问题,提出MASS算法:其通过密集自编码器学习低维主流形坐标以完成粗粒度语义分组,再在每组内利用TopK稀疏自编码器执行感知质量的稀疏特征覆盖。在Vision Flan与LLaVA-CoT上的实验表明,MASS在多种预算下均优于主流数据选择基线,且在若干设置中仅用小部分数据即可达到或超过全量数据训练的效果。

英文摘要

As supervised fine-tuning data continues to scale, selecting high-value subsets from large candidate pools is crucial for reducing training cost and improving model performance. Existing methods often measure diversity directly in the original embedding space, where geometric metrics entangle dominant semantic directions, fine-grained supervision differences, and local noise. We address this limitation by formulating data selection as a coarse-to-fine hierarchical coverage problem and propose MASS. MASS learns low-dimensional principal manifold coordinates with a dense autoencoder for coarse semantic grouping, and then performs quality-aware sparse feature coverage within each group using a TopK sparse autoencoder. Experiments on Vision Flan and LLaVA-CoT show that MASS consistently outperforms strong data selection baselines across multiple budgets, and in several settings matches or surpasses full data training with only a small subset of data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑