arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31705cs.CVcs.CL

MDL校准的显著性增益对编码:用于子词标记化的复制感知自动停止

MDL-Calibrated Significance-Gain Pair Encoding: Replication-Aware Automatic Stopping for Subword Tokenization

Azam Nouri

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出MDL-SG,一种三阶段子词标记化方法,通过显著性增益排序、复制检验和MDL效用评估实现自动停止,在WikiText-103和TinyGPT实验中优于SG-BPE和频率BPE,取得更低BPC。

中文摘要 AI 辅助

字节对编码(BPE)通过贪婪的成对合并构建子词词汇表,但传统BPE需要外部指定合并次数或目标词汇表大小。显著性增益对编码(SG-BPE)用基于观察到的配对在独立模型下超过其预期共现强度的统计标准取代了仅基于频率的选择。本文介绍了MDL校准的显著性增益对编码(MDL-SG),这是一个三阶段过程,分别处理发现、复制和效用。候选对在发现分区上按显著性增益排序,在独立分区上使用精确的单侧超几何检验(每次迭代进行Benjamini-Hochberg校正)测试其复制性,然后在效用分区上使用最小描述长度(MDL)标准进行评估。当没有复制的候选产生正的留出MDL增益时,合并自动停止。在WikiText-103上,对于120K、250K和500K字符的标记化训练样本,MDL-SG分别在209、433和847次合并处停止。在500K字符时,它选择存储词汇表为1,017个标记,而无需预先指定词汇表大小。在计算匹配的TinyGPT实验中,使用相同的2,024,448参数模型和每个语言模型500次优化器更新,MDL-SG实现了验证/测试BPC为3.2612/3.2436,而SG-BPE的测试BPC为3.2894,频率BPE为3.3493。频率BPE实现了更强的原始压缩,而MDL-SG实现了更低的BPC,表明面向压缩的合并选择与语言模型效用不必一致。

英文摘要

Byte-Pair Encoding (BPE) constructs subword vocabularies through greedy pair merging, but conventional BPE requires the number of merges or target vocabulary size to be specified externally. Significance-Gain Pair Encoding (SG-BPE) replaces frequency-only selection with a statistical criterion based on how strongly an observed pair exceeds its expected co-occurrence under an independence model. This paper introduces MDL-Calibrated Significance-Gain Pair Encoding (MDL-SG), a three-stage procedure separating discovery, replication, and utility. Candidate pairs are ranked by Significance-Gain on a discovery partition, tested for replication on a separate partition using an exact one-sided hypergeometric test with per-iteration Benjamini-Hochberg correction, and then evaluated on a utility partition using a Minimum Description Length (MDL) criterion. Merging stops automatically when no replicated candidate yields positive held-out MDL gain. On WikiText-103, MDL-SG stops at 209, 433, and 847 merges for 120K, 250K, and 500K-character tokenizer-training samples, respectively. At 500K characters, it selects a stored vocabulary of 1,017 tokens without prescribing the vocabulary size in advance. In a compute-matched TinyGPT experiment with identical 2,024,448-parameter models and 500 optimizer updates per language model, MDL-SG achieves validation/test BPC of 3.2612/3.2436, compared with test BPC of 3.2894 for SG-BPE and 3.3493 for frequency BPE. Frequency BPE achieves stronger raw compression, while MDL-SG achieves lower BPC, showing that compression-oriented merge selection and language-model utility need not coincide.

补充信息

↑