arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03089cs.CL

面向大语言模型预训练的可扩展、感知频率与长度的子文档去重

Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining

Hai Wang, Chenhao Wang, Qifeng Cai, Yixiu Liu, Miao Peng, Nuo Chen, Yuanlin Tu, Chengcheng Xu, Feng Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型预训练的子文档去重难题,提出解耦重复检测与副本保留的可扩展框架,在基准数据集上实现最优模型性能,凸显显式副本保留控制的价值。

中文摘要 AI 辅助

大规模预训练语料库包含大量重复内容。尽管文档级去重已被广泛应用,但移除子文档级冗余仍具挑战性。在语料库规模下,基于后缀数组的方法通常在数据分片内独立应用,导致跨分片重复内容无法被检测,且保留行为对分片配置敏感。基于哈希的方法可实现全局精确重复计数,但通常依赖固定的副本保留策略,无法适配异构重复模式。我们提出一种可扩展的子文档去重框架,将重复检测与副本保留解耦。该框架通过自然边界分割、归一化精确哈希及分布式聚合识别重复组,随后应用显式的感知频率与长度的保留策略,为每个组分配自适应副本预算:保留更多低频或短重复内容的副本,同时更积极地删除高频或长重复内容。在FineWeb-Edu及含代码的网络语料库上开展的实验显示,经我们方法处理的数据训练出的模型,在所有评估设置中实现了最佳整体性能。这些结果凸显了显式副本保留控制的重要性。

英文摘要

Large-scale pretraining corpora contain substantial duplicate content. Although document-level deduplication is widely used, removing subdocument-level redundancy remains challenging. At corpus scale, suffix-array-based methods are commonly applied independently within shards, leaving cross-shard duplicates undetected and making the resulting retention behavior sensitive to the sharding configuration. Hash-based methods enable global exact duplicate counting, but often rely on fixed copy-retention policies that cannot accommodate heterogeneous repetition patterns. We propose a scalable subdocument deduplication framework that decouples duplicate detection from copy retention. It identifies duplicate groups through natural-boundary segmentation, normalized exact hashing, and distributed aggregation, and then applies an explicit frequency- and length-aware retention policy that allocates an adaptive copy budget to each group, retaining more copies of low-frequency or short repetitions while more aggressively deleting high-frequency or long ones. Experiments on FineWeb-Edu and a code-containing web corpus show that models trained on data processed by our method achieve the best overall performance among the evaluated settings. These results underscore the importance of explicit copy-retention control.

↑