arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

动态规划引导的分层BPE与词汇剪枝实证分析

Dynamic-Programming-Guided Hierarchical BPE and Empirical Analysis of Vocabulary Pruning

Kenny Shao

arXiv 2609.06898首次发表:更新:

AI 中文总结

提出DH-BPE方法,结合动态规划与分层依赖进行词汇剪枝,在固定词汇预算下提升压缩率,并在多个规模上优于标准BPE等基线。

AI 中文摘要

字节对编码(BPE)通过贪心对合并构建词汇表,但由此产生的合并顺序不一定能为压缩分配一个最优的固定模型可见词汇表。我们提出动态规划引导的分层BPE(DH-BPE),一种词汇构建方法,它将精确最小分词下的token暴露与BPE训练产生的层次依赖相结合。从适度超调的BPE候选词汇表开始,DH-BPE使用动态规划来衡量候选效用,并应用暴露引导、依赖感知的剪枝来选择固定大小的模型可见词汇表。我们在12K和16K目标词汇大小的主要评估中将DH-BPE与标准BPE及最近的词汇优化基线(包括Pruned BPE、MinGram和MinGram-PP)进行比较,并额外在18K下仅与MinGram进行评估。在主要的12K和16K比较中,DH-BPE在共享的精确最小token DP编码器下,相对于标准BPE、Pruned BPE和MinGram持续提高了总体压缩率。MinGram-PP在主要比较中实现了更强的总体压缩,但在跨语料库评估中,DH-BPE在超调因子f=2.0和f=3.0时优于它;在12K下,MinGram-PP仅在f=4.0和f=5.0的更大候选池下逆转了这一排序。定性分析进一步表明,DH-BPE在更晚、更完整的BPE合并与可复用的子词组件之间取得了平衡,为在固定模型可见词汇预算下改进词汇分配提供了一种实用方法。

英文摘要

Byte Pair Encoding (BPE) constructs vocabularies through greedy pair merging, but the resulting merge order does not necessarily allocate a fixed model-visible vocabulary optimally for compression. We propose Dynamic-Programming-Guided Hierarchical BPE (DH-BPE), a vocabulary-construction method that combines token exposure under exact minimum-token segmentation with the hierarchical dependencies induced by BPE training. Starting from a modestly overshot BPE candidate vocabulary, DH-BPE uses dynamic programming to measure candidate utility and applies exposure-guided, dependency-aware pruning to select a fixed-size model-visible vocabulary. We compare DH-BPE against Standard BPE and recent vocabulary-optimization baselines, including Pruned BPE, MinGram, and MinGram-PP, in primary evaluations at 12K and 16K target vocabulary sizes, with an additional 18K evaluation against MinGram only. Across the primary 12K and 16K comparisons, DH-BPE consistently improves aggregate compression over Standard BPE, Pruned BPE, and MinGram under a shared exact minimum-token DP encoder. MinGram-PP achieves stronger aggregate compression in the primary comparisons, but DH-BPE outperforms it at overshoot factors f = 2.0 and f = 3.0 in cross-corpus evaluation; at 12K, MinGram-PP reverses this ordering only with the substantially larger candidate pools at f = 4.0 and f = 5.0. Qualitative analysis further shows that DH-BPE balances later, more complete BPE merges with reusable subword components, providing a practical approach to improving vocabulary allocation under a fixed model-visible vocabulary budget.

Comments21 pages, 5 figures, 3 tables, and 1 algorithm

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑