arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向语言模型预训练的可扩展数据多样化:基于杠杆分数采样

Towards Scalable Data Diversification for Language Model Pretraining via Leverage Score Sampling

Zailin Ma, Quzhe Huang, Yujun Li, Congyuan Rao, Yaodong Yang

arXiv 2609.32484首次发表:更新:

AI 中文总结

针对预训练数据选择中质量与多样性的矛盾,提出杠杆分数采样方法,以高效迭代扩展数据行列式体积,实现72倍加速、多样性提升9.2%,并在下游任务与代码数据上取得更优性能。

AI 中文摘要

语言模型预训练的数据选择面临着质量与多样性之间的根本性张力。虽然质量过滤在经验上有效,但常常导致多样性崩溃:通过偏好与高质量参考语料库(如教育或问答风格数据)相似的文本,它系统性地排除了来自代表性不足领域的宝贵数据。相比之下,多样化选择保持了领域平衡并促进了稳健的下游性能,然而现有方法要么侧重于间接增强多样性的覆盖导向目标,要么通过昂贵的协方差矩阵重计算直接优化多样性,这限制了可扩展性。为解决这些问题,我们引入了杠杆分数采样(Leverage Score Sampling,简称Lev),该方法通过杠杆分数(一种计算高效的标准,消除了矩阵重计算并实现了可扩展选择)迭代地选择那些能最大程度扩展嵌入数据行列式体积的样本。实验上,Lev实现了高达72倍的加速,并将数据集多样性(以Vendi分数衡量)较强大的多样化基线DiSF提升了9.2%。在CommonCrawl(CC)网络数据选择中,Lev在七个下游任务上的准确率较现有基线最高提升了1.31%。对于难以明确定义稳健质量标准的领域(如代码),Lev作为一种有效的无监督筛选替代方案:在StarCoderData上,所选子集将每字节比特数(bits-per-byte)较DiSF降低了3.08%。值得注意的是,我们发现了质量过滤的跨领域崩溃现象:经DCLM-fastText过滤的CC数据未能保留足够的代码相关内容,导致代码性能劣于Lev选择的数据。这些发现主张将多样性感知实践整合到质量过滤中,以实现语言模型预训练中更有效的数据筛选。

英文摘要

Data selection for language model pretraining faces a fundamental tension between quality and diversity. While quality filtering is empirically effective, it often induces diversity collapse: by favoring texts similar to high-quality reference corpora (e.g., educational or QA-style data), it systematically excludes valuable data from underrepresented domains. In contrast, diversified selection preserves domain balance and encourages robust downstream performance, yet existing methods either focus on coverage-oriented objectives that indirectly enhance diversity, or directly optimize for diversity via costly covariance matrix recomputation that limits scalability. To address these issues, we introduce \textbf{Leverage Score Sampling (Lev)}, which iteratively selects samples that maximally expand the determinantal volume of the embedded data via leverage scores, a computationally efficient criterion that eliminates matrix recomputation and enables scalable selection. Empirically, Lev delivers up to $72\times$ speedup and improves dataset diversity, measured by the Vendi score, by $9.2\%$ over the strong diversification baseline \textbf{DiSF}. On CommonCrawl (CC) web data selection, Lev improves accuracy across seven downstream tasks by up to $1.31\%$ over existing baselines. For domains where robust quality criteria are inherently difficult to define (e.g., code), Lev serves as an effective unsupervised curation alternative: on StarCoderData, the selected subset reduces bits-per-byte by $3.08\%$ over DiSF. Notably, we uncover a cross-domain collapse of quality filtering: CC data filtered by DCLM-fastText fail to retain sufficient code-related content, yielding inferior code performance relative to Lev-selected data. These findings advocate for integrating diversity-aware practices into quality filtering for more effective data curation in language model pretraining.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑