超越每个字母两字节:西里尔字母AI系统中的分词开销
Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems
AI总结:
该研究量化了多语言分词器对西里尔字母语言的分词开销,评估了LLMLingua-2、平衡字节级BPE等缓解策略,指出训练数据分配是开销成因且可在两阶段缓解。
AI中文摘要:
现代多语言分词器对乌克兰语及其他代表性不足的西里尔字母语言的分词拆分程度通常远高于英语,造成成本与上下文容量的差异。我们在9种生产级分词器和5种具有标准化西里尔与拉丁字母表示的语言(覆盖837万种词形)中量化了该开销。在语料库基准测试中,乌克兰语在现代分词器上的分词开销为68%-121%,在旧版cl100k分词器上达220%,通过BrUK和Brown语料库的全文生成能力衡量。在经独立验证的英语基准子集上,开销与西里尔词汇分配呈负相关,但相关性无统计学意义(斯皮尔曼相关系数rho=-0.536,p=0.215,样本量n=7)。我们评估了两种缓解策略:在包含1536种产品和145个查询的电商RAG基准测试中,LLMLingua-2将乌克兰语输入长度缩短了47%-49%,且在80个可检索案例中未出现压缩导致的价值损失;采用20万词汇上限训练的平衡字节级BPE分词器(实际收敛至158184个条目),将保留的乌克兰语/英语(UK/EN)比例从2.22倍降至1.30倍。在大多数分词器上,罗马化会使乌克兰语的分词数量增加2%-19%。在五种语言中,分词效率偏向网络数据中更普遍的字母体系。这些发现表明,训练数据分配是西里尔字母分词开销的成因,且在推理阶段和分词器设计阶段均存在缓解可能。
英文摘要:
Modern multilingual tokenizers often fragment Ukrainian and other underrepresented Cyrillic-script languages more heavily than English, creating disparities in cost and context capacity. We quantify this overhead across nine production tokenizers and five languages with standardized Cyrillic and Latin representations, covering 8.37 million word forms. On a corpus benchmark, Ukrainian shows 68-121% token overhead on modern tokenizers and 220% on the older cl100k, measured through full-text fertility on the BrUK and Brown corpora. Overhead is negatively associated with Cyrillic vocabulary allocation in the subset with independently verified English baselines, although the association is not statistically significant (Spearman rho = -0.536, p = 0.215, n = 7). We evaluate two mitigation strategies. LLMLingua-2 reduces Ukrainian input length by 47-49% on an e-commerce RAG benchmark of 1,536 products and 145 queries, with no compression-induced value losses among 80 retrievable cases. A balanced byte-level BPE tokenizer trained with a 200K vocabulary cap, converging at 158,184 actual entries, reduces the held-out UK/EN ratio from 2.22x to 1.30x. Romanization increases Ukrainian token counts by 2-19% on most tokenizers. Across the five languages, tokenization efficiency favors the script more prevalent in web data. These findings indicate that training data allocation contributes to Cyrillic tokenization overhead and that mitigation is possible at both inference and tokenizer-design stages.