发表机构
Tsinghua University; SiliconProspect AI; Nanyang Technological University; Alibaba Group(清华大学; 硅景人工智能; 南洋理工大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对中文网页污染审计的三项挑战,提出Sampled-BPE轻量级审计流程,在大幅降低运行时间和内存消耗的同时保持精度,审计多类语料库并发布分层中文网页词元数据集。
AI 中文摘要
中文网页污染已在大语言模型(LLMs)中显现,这推动了对上游中文语料库的审计工作。然而,审计此类语料库面临三项挑战:其一,其网页级规模使得全量扫描成本高昂;其二,过往分析通常过于粗糙,无法揭示词元层面的污染;其三,中文网页污染具有隐蔽性且变化迅速。我们提出Sampled-BPE,这是一种轻量级词元级审计流程,通过采样小部分数据并训练BPE分词器来识别污染词元。实验表明,Sampled-BPE在保持可用估计精度的同时,大幅降低了运行时间和内存消耗:速度提升148.4倍,内存减少35.8倍,而污染类别的相对误差仅为4.25%。我们将该流程应用于11个开源中文语料库及2021至2026年的6个中文Common Crawl快照。审计结果显示,开源语料库中存在普遍但分布不均的污染,以及高度污染且随时间变化的中文网页内容。我们还发布了一个分层中文网页词元数据集,包含66万余个词元记录,每个记录均具备网页上下文、类别和解释字段,以树形结构组织,支持对污染的审查与溯源。
英文摘要
Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three challenges: (1) their web-scale size makes full scan costly; (2) prior analyses are often too coarse to expose token-level pollution; (3) Chinese web pollution is implicit and rapidly changing. We propose Sampled-BPE, a lightweight token-level auditing pipeline that sample a small subset and train BPE tokenizer to surface polluted tokens. Experiments show that Sampled-BPE preserves usable estimates while substantially reducing runtime and memory: a 148.4 $\times$ speedup and a 35.8 $\times$ memory reduction induce only 4.25% relative error for pollution categories. We apply the pipeline to 11 open Chinese corpora and 6 Chinese Common Crawl snapshots from 2021 to 2026. The audit reveals widespread but uneven pollution across open corpora, as well as highly polluted and temporally shifting Chinese web content. We further release a hierarchical Chinese web token dataset with 660k+ token records, each with web context, category, and explanation fields, organized as trees to support review and tracing of pollution.