arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

元音符号并非字母:多语言分词器生成能力的预分词上限

Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility

Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun

arXiv 2608.26449首次发表:更新:

发表机构

Karela Technologies Inc.(卡雷拉科技公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出多语言分词器因错误拆分元音附标文字的元音符号导致生成能力受限,提出修复方案并通过实验验证其能提升尼泊尔语模型性能,还统计了相关分词器的部署情况。

AI 中文摘要

使用HuggingFace ByteLevel预分词器的字节级BPE分词器继承了GPT-2的单词正则表达式,其中单词被定义为\p{L}+,即一个或多个Unicode字母。在元音附标文字中,元音以组合符号形式书写;因此该模式会在每个元音符号处拆分单词。由于BPE仅在预分词内部进行合并,无论词汇量或语料库组成如何,这些拆分都会在训练过程中持续存在。我们将这种效应形式化为生成能力的无训练下限。在来自平行语料库的26种语言中,所有17种元音附标文字均受影响,范围从1.47倍(藏语)到9.02倍(泰语),而拉丁、西里尔、谚文和汉字则显示恰好1.00倍。对于5种仅在该字符类上不同的匹配分词器对,其预测下限的偏差在2.2%以内,尼泊尔语上的单词分词数分别为4.78和1.58。当训练语料库中尼泊尔语的占比从5%提升至95%时,存在缺陷的分词器几乎没有变化(仅1.7%),而修复后的分词器则变化了33.9%,无需检查任何代码即可区分结构上限与数据短缺。我们训练了3个仅分词器不同的2.68亿参数模型;在计算量相同的情况下,修复后的变体在保留的尼泊尔语位每字节上降低了4.43%,在相同字节数下使用1.59倍计算量时仍保持领先。对3479个HuggingFace仓库的统计发现,63.3%的下载量最高的文本生成模型使用了仅包含字母的单词类,占这些模型总下载量的72.5%。GPT-4o的o200k模式已使用感知符号的单词类,因此该修复本身属于现有技术。我们量化了其价值,展示了如何仅通过症状识别其缺失,绘制了其适用的文字系统,测量了其部署范围,并发布了一个包含65536个条目的尼泊尔语-英语分词器及配套工具,可在笔记本电脑上从公共数据复现此处的所有数值。

英文摘要

Byte-level BPE tokenizers that use the HuggingFace ByteLevel pre-tokenizer inherit GPT-2's word regex, where a word is defined as \p{L}+, one or more Unicode letters. In abugida scripts, vowels are written as combining marks; this pattern therefore splits each word at every vowel sign. Since BPE merges only within a pre-token, those splits persist through training regardless of vocabulary size or corpus composition. We formalise this effect as a training-free lower bound on fertility. Across 26 languages from a parallel corpus, every one of the 17 abugidas is affected, ranging from 1.47x (Tibetan) to 9.02x (Thai), whereas Latin, Cyrillic, Hangul, and Han show exactly 1.00x. For 5 languages, matched tokenizer pairs that differ only in this character class fall within 2.2% of the predicted floor, scoring 4.78 versus 1.58 tokens per word on Nepali. When the Nepali share of the training corpus is swept from 5% to 95%, the broken tokenizer barely shifts at all (1.7%) while the fixed one shifts 33.9%, which separates a structural ceiling from a data shortage without needing to inspect any code. We train three 268M models that differ only in their tokenizer; the fixed variant achieves 4.43% lower held-out Nepali bits per byte at equal compute, and it still leads when given the same bytes with 1.59x the compute. A census of 3,479 HuggingFace repositories finds the letters-only word class present in 63.3% of the most-downloaded text-generation models, accounting for 72.5% of their downloads. GPT-4o's o200k pattern already uses a mark-aware word class, making the repair itself prior art. We quantify its value, show how to recognise its absence from symptoms alone, map which scripts it reaches, measure how widely it is deployed, and release a 65,536-entry Nepali-English tokenizer with a harness that regenerates every number here from public data on a laptop.

Comments14 pages, 2 figures, 12 tables. Code, tokenizer and reproduction harness: https://github.com/sajalregmi/arkios-tokenizer and https://huggingface.co/sajalregmi4/arkios-tokenizer

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑