arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01984cs.CLcs.LG

通用字节级编码:UTF-8/UTF-16路由以减少跨文字令牌预算差异

Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities

Hyunsik Kim, Youngmoon Jung

首次发表
浏览论文内容

中文总结 AI 辅助

提出通用字节级编码(UBE),通过UTF-8/UTF-16路由降低多语言令牌预算差异,保持解码精确,匹配语言模型质量并提升上下文效率。

中文摘要 AI 辅助

字节级字节对编码(BBPE)分词器对多语言大语言模型(LLM)具有吸引力,因为它们覆盖所有Unicode文本。然而,在基于UTF-8的BBPE中,许多文字从比英语更高的回退成本开始:当无法应用学习到的合并时,多字节字符需要多个字节派生的符号。我们将这种最坏情况下的合并前成本称为编码下限。较高的下限会增加令牌数量、每次请求的成本,并缩小可用上下文。改变文本编码可以减少这种差距,但单一的全局编码可能使混合文字文本中已经高效的英语片段变得更加昂贵。我们提出通用字节级编码(UBE),一种双字母表分词器,将1-2字节的UTF-8字符保留在UTF-8路径上,同时将3-4字节的UTF-8字符路由到UTF-16。这降低了具有高令牌溢价(相对于英语的令牌数量)的文字中3字节基本多文种平面(BMP)字符的编码下限,而不会提高混合文字文本中已经高效的片段的下限。UBE仅改变呈现给字节对编码(BPE)的字节表示;合并规则保持标准,精确解码得以保留。UBE还可与替代边界策略和基于形态的表示组合使用。在Unicode 17审计中,UBE精确往返所有Unicode标量值以及官方规范化、字素分割和表情符号测试套件中的所有输入。在内在评估中,UBE降低了英语归一化令牌计数比率的离散度,减少了跨语言令牌预算差异。在多语言语言模型(LM)实验中,UBE匹配BBPE的LM质量。在主要的多语言设置中,UBE对高溢价文字减少的令牌数量最多,并略微降低英语令牌数量,在固定令牌预算下提供更多可用上下文,并在内容匹配基准中加快提示处理。

英文摘要

Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 characters on the UTF-8 path while routing 3-4-byte UTF-8 characters through UTF-16. This lowers the encoding floor for 3-byte Basic Multilingual Plane (BMP) characters in scripts with high token premiums (token counts relative to English) without raising it for already-efficient spans in mixed-script text. UBE changes only the byte representation presented to byte-pair encoding (BPE); the merge rule remains standard, and exact decoding is preserved. UBE also composes with alternative boundary policies and morphology-based representations. In a Unicode 17 audit, UBE exactly round-trips all Unicode scalar values and all inputs in the official normalization, grapheme-break, and emoji test suites. Across intrinsic evaluations, UBE lowers dispersion in English-normalized token-count ratios, reducing cross-lingual token-budget disparity. In multilingual language model (LM) experiments, UBE matches BBPE's LM quality. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and slightly lowers English token counts, yielding more usable context under fixed token budgets and faster prompt processing in content-matched benchmarks.

发表机构

  • Samsung Research(三星研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑