arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分词器代价:量化并解释印度语言子词分词的跨语言成本

The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages

Priyansh Srivastava

arXiv 2607.24276首次发表:更新:

发表机构

Sirena Ai India(印度Sirena人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究量化印度语言子词分词跨语言成本,测量六种分词器在十四种语言上的分词丰富度,揭示字节对合并失败是高代价主因,多语言分词器可降低代价,还发现分词与阅读理解性能关联受语言资源可用性影响。

AI 中文摘要

大型语言模型通过子词分词器处理文本,由于这些分词器主要在以英语为中心的语料库上训练,对许多非英语语言存在系统性且常被忽视的劣势。本文利用FLORES - 200平行语料库量化印度语言的分词器代价,测量六种分词器和十四种语言的分词丰富度。在cl100k_base下,印度语言相对于英语平均有8.0倍的分词代价,如马拉雅拉姆语达13.0倍。确定主要机制是字节对合并失败致文本碎片化,多语言分词器可降低代价。还量化了实际影响,发现分词丰富度与阅读理解性能的关联很大程度由语言资源可用性决定。

英文摘要

Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are trained predominantly on English-centric corpora, they introduce a systematic and often overlooked disadvantage for many non-English languages. In this work, we quantify this tokenizer tax for Indian languages using the FLORES-200 parallel corpus, measuring tokenization fertility across six widely used tokenizers and fourteen languages. Under cl100k_base (used by GPT-3.5 and GPT-4), Indian languages experience an average 8.0x tokenization tax relative to English, reaching 13.0x for Malayalam, reducing the effective context window to as little as 12% of that available to English users for equivalent semantic content. We identify the primary mechanism behind this disparity: failed byte-pair merges that leave text fragmented into single-byte tokens, with merge failure strongly correlating with tokenizer tax (Pearson r = 0.89). We further show that this phenomenon is not an inherent property of Indic scripts but a consequence of tokenizer design. Multilingual tokenizers such as XLM-R and OpenAI's o200k_base reduce the average Indic tokenizer tax by 73%, demonstrating that the disparity is largely remediable. Beyond token statistics, we quantify a practical consequence by showing that, under fixed context budgets, Indian-language documents preserve substantially less original content than equivalent English documents. Finally, we examine the relationship between tokenizer fertility and reading comprehension performance on the Belebele benchmark, finding that the apparent correlation is largely explained by language resource availability rather than tokenizer behavior alone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑