面向大语言模型预训练的可扩展、感知频率与长度的子文档去重
Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining
浏览论文内容
中文总结 AI 辅助
针对大语言模型预训练的子文档去重难题,提出解耦重复检测与副本保留的可扩展框架,在基准数据集上实现最优模型性能,凸显显式副本保留控制的价值。
中文摘要 AI 辅助
大规模预训练语料库包含大量重复内容。尽管文档级去重已被广泛应用,但移除子文档级冗余仍具挑战性。在语料库规模下,基于后缀数组的方法通常在数据分片内独立应用,导致跨分片重复内容无法被检测,且保留行为对分片配置敏感。基于哈希的方法可实现全局精确重复计数,但通常依赖固定的副本保留策略,无法适配异构重复模式。我们提出一种可扩展的子文档去重框架,将重复检测与副本保留解耦。该框架通过自然边界分割、归一化精确哈希及分布式聚合识别重复组,随后应用显式的感知频率与长度的保留策略,为每个组分配自适应副本预算:保留更多低频或短重复内容的副本,同时更积极地删除高频或长重复内容。在FineWeb-Edu及含代码的网络语料库上开展的实验显示,经我们方法处理的数据训练出的模型,在所有评估设置中实现了最佳整体性能。这些结果凸显了显式副本保留控制的重要性。
英文摘要
Large-scale pretraining corpora contain substantial duplicate content. Although document-level deduplication is widely used, removing subdocument-level redundancy remains challenging. At corpus scale, suffix-array-based methods are commonly applied independently within shards, leaving cross-shard duplicates undetected and making the resulting retention behavior sensitive to the sharding configuration. Hash-based methods enable global exact duplicate counting, but often rely on fixed copy-retention policies that cannot accommodate heterogeneous repetition patterns. We propose a scalable subdocument deduplication framework that decouples duplicate detection from copy retention. It identifies duplicate groups through natural-boundary segmentation, normalized exact hashing, and distributed aggregation, and then applies an explicit frequency- and length-aware retention policy that allocates an adaptive copy budget to each group, retaining more copies of low-frequency or short repetitions while more aggressively deleting high-frequency or long ones. Experiments on FineWeb-Edu and a code-containing web corpus show that models trained on data processed by our method achieve the best overall performance among the evaluated settings. These results underscore the importance of explicit copy-retention control.