arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

已发布的大语言模型词汇表能否支持隐藏语料库的词元级估计?

Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?

Qingjie Zhang, Xingzhang Ren, Zixuan Chen, Jinfeng Li, YueFeng Chen, Yitong Yang, Hui Xue, Dayiheng Liu, Han Qiu

arXiv 2608.10690首次发表:更新:

发表机构

Tsinghua University; Alibaba Group(清华大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对隐藏语料库的词元级估计问题,提出分位数引导密度估计(QGDE)方法,利用已发布的LLM分词器词汇表实现了低误差的细粒度语料估计,为相关任务提供了新的有效信号来源。

AI 中文摘要

预训练语料库的构成会影响大语言模型(LLM)的能力,但即便模型权重已发布,语料库构成往往仍处于隐藏状态。过往研究曾从已发布的分词器词汇表中推断语料混合比例或追踪特定词元组;与之不同,我们针对任意目标词元估计其对应的语料比例。我们首先证明,在不同语料上训练得到的字节对编码(BPE)分词器具有稳定的词元ID-比例分布,这一发现为从已知语料向在隐藏语料上训练得到的目标分词器进行分布迁移提供了依据。随后,我们提出分位数引导密度估计(QGDE)方法,该方法通过多个分位数趋势近似上述分布,并结合局部密度加权生成词元级估计结果。在受控设置及使用已发布SmolLM分词器的实际场景中,QGDE在词元级估计任务上的平均相对误差低至3.00%,聚合为类别级语料混合比例后误差为3.08%。上述结果表明,已发布的分词器词汇表除了可用于粗粒度的语料构成推断外,还能为细粒度的语料估计提供有用信号。

英文摘要

Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus mixtures or traced specific token groups from released tokenizer vocabularies; in contrast, we estimate corpus ratios for arbitrary target tokens. We first show that BPE tokenizers trained on different corpora share stable token ID--ratio distributions, motivating distribution transfer from known corpora to a target tokenizer trained on hidden corpora. We then propose Quantile-Guided Density Estimation (QGDE), which approximates this distribution with multiple quantile trends and uses local density weighting to produce token-level estimates. In controlled settings and a realistic setting using the released SmolLM tokenizer, QGDE achieves mean relative errors as low as 3.00% for token-level estimation and 3.08% after aggregation into category-level mixtures. These results suggest that released tokenizer vocabularies provide a useful signal for fine-grained corpus estimation beyond coarse composition inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑