arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

贪心最长匹配分词的联合优化

Joint Optimization for Greedy Longest-match Tokenization

Adhiraj Singh, Deepanshu Mody, Ghina Al Shdaifat, Hamza Alshamy, Adam Wiemerslage, Varshini Reddy, Craig W. Schmidt

arXiv 2607.23362首次发表:更新:

发表机构

Center for Data Science, New York University; Kensho Technologies(纽约大学数据科学中心; 肯绍科技公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对贪心最长匹配解码,提出联合优化方法JOLT,将词汇学习设为整数规划,通过贪心一致性约束等确保训练与部署时分词对齐,求解线性规划松弛扩大优化规模,实验表明其能缩小与最优压缩差距,减少令牌数量。

AI 中文摘要

近期研究表明,子词词汇表可经训练以针对特定推理规则优化压缩,而非依赖如字节对编码(BPE)等贪心启发式方法。本文将此方法扩展至贪心从左到右最长匹配解码,这是WordPiece基础的快速且广泛使用的推理规则。引入贪心最长匹配分词联合优化(JOLT),将词汇学习表述为词汇选择和分割选择变量上的整数规划。贪心一致性约束确保每个优化分割与所选词汇表下最长匹配解码产生的分割完全匹配,使训练目标与部署时的分词对齐。为扩大优化规模,求解线性规划松弛并仅为未解决的预令牌选择性引入高阶分割。结果松弛几乎是整数:舍入解在训练范围内的线性规划下界的0.008 - 0.176%内。该界限还表明BPE在贪心最长匹配解码下已处于最佳可实现压缩的1 - 2%内,而JOLT缩小了剩余差距的89.6 - 99.4%。在四个训练范围以及32000和64000词汇量的留出验证数据上,JOLT比BPE产生的令牌少0.78%,且随着训练范围增加改进通常增大。这些结果表明推理对齐的词汇优化可恢复BPE留下的大部分有限压缩空间,同时提供接近最优的证明。

英文摘要

Recent work has shown that subword vocabularies can be trained to optimize compression for a specific inference rule rather than relying on greedy heuristics such as Byte Pair Encoding (BPE). We extend this approach to greedy left-to-right longest-match decoding, the fast and widely used inference rule underlying WordPiece. We introduce Joint Optimization for Greedy Longest-Match Tokenization (JOLT), which formulates vocabulary learning as an integer program over vocabulary-selection and segmentation-choice variables. Greedy-consistency constraints ensure that each optimized segmentation exactly matches the segmentation produced by longest-match decoding under the selected vocabulary, aligning the training objective with deployment-time tokenization. To scale the optimization, we solve a linear programming relaxation and selectively introduce higher-order segmentations only for unresolved pretokens. The resulting relaxation is nearly integral: rounded solutions fall within 0.008 - 0.176 % of the LP lower bound on the training scope. The bound also shows that BPE is already within 1 - 2 % of the best achievable compression under greedy longest-match decoding, while JOLT closes 89.6 - 99.4 % of the remaining gap. On held-out validation data across four training scopes and vocabulary sizes of 32,000 and 64,000, JOLT produces up to 0.78 % fewer tokens than BPE, with improvements generally increasing as the training scope grows. These results demonstrate that inference-aligned vocabulary optimization can recover most of the limited compression headroom left by BPE while providing a certificate of near-optimality.

Comments18 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑