arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PUFFER:面向持续演化语料库的增量模糊去重算法

PUFFER: Incremental Fuzzy Deduplication for Continuously Evolving Corpora

Xiao Yang, Erik Edward Aldape, Beren Millidge

arXiv 2608.28622首次发表:更新:

发表机构

Zyphra(齐普拉)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 PUFFER 增量模糊去重算法,通过 LSH 带存储与 T-扇出分层压缩实现低内存、低成本的持续演化语料库去重,在百亿级文档场景下比同类工具快数十倍,已应用于超 300 亿文档并开源。

AI 中文摘要

大语言模型训练语料库通过连续发布不断增长,且常包含冗余内容,因此每次发布都需与自身及累计历史数据进行去重。在万亿 token 规模下,这要求增量式数据摄入、受限驻留内存、确定性重试以及数据集级生命周期控制,且无需重复构建全量语料库。我们提出 PUFFER(Provenance-aware Updatable Fuzzy Filtering for Evolving Repositories,面向演化仓库的溯源感知可更新模糊过滤算法),一种基于 MinHash-LSH 的模糊去重流水线,其核心设计包含两点:其一,PUFFER 将每个 LSH 带存储为不可变、带数据集标签、内存映射的排序段,可实现精确的历史带键成员资格检查,且内存占用不随语料库规模成比例增长;其二,T-扇出分层压缩机制会定期合并段以控制筛选扇出,通过牺牲更低的查询成本换取额外的索引维护写入,同时保留成员资格判定。在 N 个摄入键和 K 个等规模发布的场景下,PUFFER 的累计维护成本为 O(N log N log_T K),而重复快照重建的成本为 Θ(KN)。带数据集标签的段还支持数据集级撤回:对于未压缩或受保护的数据集,撤回操作耗时为常数;而压缩后的撤回操作仅需重构受影响的合并段,即使原始数据集不可用也可完成。在我们的实现中,PUFFER 在单进程下约 1.75 小时完成了 10 亿文档的累计索引阶段摄入,16 带索引下每个文档仅占用 128 字节内存;而经典驻留式 MinHash-LSH 表每个文档约占用 6.5 KB 内存,超出 900 GiB 的内存上限。在以 10 亿文档为上限的 10 小时对比实验中,PUFFER 比 LSHBloom 快 11 倍,比 Milvus-LSH 快 35 倍。目前 PUFFER 已应用于超过 300 亿文档,我们将其作为开源软件发布,链接为 this https URL。

英文摘要

Large language model training corpora grow through successive, often redundant releases, so each release must be deduplicated against both itself and the accumulated history. At trillion-token scale, this requires incremental ingestion, bounded resident memory, deterministic retry, and dataset-scoped lifecycle control without repeated corpus-wide rebuilding. We introduce PUFFER (Provenance-aware Updatable Fuzzy Filtering for Evolving Repositories), a MinHash-LSH fuzzy-deduplication pipeline built around two design choices. First, PUFFER stores each LSH band as immutable, dataset-tagged, memory-mapped sorted segments, enabling exact historical band-key membership checks without RAM proportional to corpus size. Second, T-fanout tiered compaction periodically merges segments to control screening fanout, trading lower query cost against additional index-maintenance writes while preserving membership decisions. Across N ingested keys and K equal-sized releases, PUFFER's cumulative maintenance cost is O(N log N log_T K), compared with Theta(KN) for repeated snapshot rebuilding. Dataset-tagged segments also support dataset-scoped withdrawal: removal is constant-time for uncompacted or protected datasets, while post-compaction withdrawal reconstructs only the affected merged segment, even if the original dataset is unavailable. In our implementation, PUFFER completed cumulative index-stage ingestion for one billion documents in about 1.75 hours in a single process, using 128 bytes per document for a 16-band index. A classical resident MinHash-LSH table required about 6.5 KB per document and exceeded a 900 GiB RAM cap. In a ten-hour comparison capped at one billion documents, PUFFER was 11x faster than LSHBloom and 35x faster than Milvus-LSH. PUFFER is deployed on more than 30 billion documents, and we release it as open-source software at https://github.com/Zyphra/puffer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑