发表机构
Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对开放网络语料库预训练数据优化的局限性,提出质量、冗余性和多样性联合优化范式,结合双轨清理与混合去重,构建CuraWeb语料库,实验证明其在多基准测试中性能显著优于基线。
AI 中文摘要
通过高度选择性过滤器精心策划的开放网络语料库,如FineWeb-Edu和DCLM,构成了大语言模型预训练数据的核心,并显著提升了大语言模型的性能。然而,这些流程通常依赖单一优化目标,不可避免地缩小了分布多样性并边缘化了长尾知识,从而限制了数据覆盖范围并未充分利用开放网络的巨大潜力。为解决这一局限性,我们提出了一种新颖的策划范式,从线性修剪转向质量、冗余性和多样性的联合优化。该框架将双轨清理(基于规则和模型驱动)与混合去重(n-gram和语义)相结合,同时采用多目标采样器来平衡信息质量和分布广度。将此框架应用于Common Crawl,我们构建了CuraWeb,一个2T-token的英语语料库。与现有资源不同,CuraWeb通过恢复更全面的数据分布、增强多样性和最小化冗余,建立了工业级数据策划标准,实现了跨不同领域长尾知识的更广泛覆盖。在3B规模的实验评估表明,CuraWeb显著优于现有基线,在广泛的基准测试中平均性能提升1.8%,特别是在知识密集型和推理任务中。
英文摘要
Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly advanced LLM performance. However, these pipelines typically rely on singular optimization objectives, which inevitably narrows distributional diversity and marginalizes long-tail knowledge, thereby restricting data coverage and underutilizing the vast potential of the open web. To address this limitation, we propose a novel curation paradigm that shifts from linear pruning to the joint optimization of quality, redundancy, and diversity. This framework synergizes dual-track cleaning (rule-based and model-driven) with hybrid deduplication (n-gram and semantic), while employing a multi-objective sampler to balance informational quality with distributional breadth. Applying this framework to Common Crawl, we construct CuraWeb, a 2T-token English corpus. Unlike existing resources, CuraWeb establishes an industrial-grade standard for data curation by recovering a more holistic data distribution with enhanced diversity and minimal redundancy, achieving broader coverage of long-tail knowledge across diverse domains. Experimental evaluations at the 3B scale demonstrate that CuraWeb significantly outperforms state-of-the-art baselines, yielding an average performance gain of 1.8\% across a wide range of benchmarks, particularly in knowledge-intensive and reasoning tasks.