arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

标记化溢价的衡量:服务不足语言社区的成本审计

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

Avijit Roy, Proma Roy, Hrishitva Patel

arXiv 2608.09046首次发表:更新:

发表机构

John Jay College of Criminal Justice, City University of New York; The City College of New York, City University of New York; University of Texas at San Antonio(纽约城市大学约翰杰刑事司法学院; 纽约城市大学城市学院; 德克萨斯大学圣安东尼奥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过构建标记化公平审计(TEA)基准,测量了孟加拉语、印地语等5种语言的标记化溢价,发现标记化会给服务不足语言社区带来经济和功能障碍,凸显需重视标记化的公平性。

AI 中文摘要

大型语言模型正日益被用作通用教育和技术辅助系统,但其底层基础设施并未平等对待各种语言。一个被研究不足的差异来源是标记化:语义等价的内容在不同语言中可能需要明显不同的标记数量,这会影响API成本、延迟,以及模型被调用前可用的上下文长度。我们引入了标记化公平审计(Tokenization Equity Audit, TEA),这是一种用于衡量技术辅导内容中标记化溢价的可复现基准。TEA在从英语翻译成孟加拉语、印地语、阿拉伯语、泰米尔语和约鲁巴语的120项Python调试语料库上评估了三种广泛使用的标记器:GPT-4o的o200k基础版、Qwen2.5-7B和Mistral-7B。孟加拉语和印地语作为主要验证案例,其余语言提供跨脚本和跨语系的探索性比较。在该语料库中,孟加拉语所需的GPT-4o标记数量是英语的1.56倍,将名义上128k标记的上下文窗口,缩减为相同语义内容下等效于英语的有效82k标记容量;使用Qwen2.5和Mistral标记器时,孟加拉语所需标记数量最多为英语的4.5倍。约鲁巴语尽管使用拉丁字母,却表现出最高的GPT-4o标记化溢价,达2.37倍,这表明标记化不公平无法仅用脚本语系来解释。这些结果表明,标记化会造成可衡量的经济和功能障碍,凸显需将标记化视为与公平相关的基础设施层,以服务不足的语言社区,尤其是在教育系统依赖低成本或离线可用AI工具的场景中。

英文摘要

Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure does not treat languages equally. One underexamined source of disparity is tokenization: semantically equivalent content can require substantially different token counts across languages, affecting API cost, latency, and usable context length before a model is invoked. We introduce the Tokenization Equity Audit (TEA), a reproducible benchmark for measuring tokenization premiums in technical tutoring content. TEA evaluates three widely used tokenizers, GPT-4o's o200k base, Qwen2.5-7B, and Mistral-7B, on a 120-item Python debugging corpus translated from English into Bengali, Hindi, Arabic, Tamil, and Yoruba. Bengali and Hindi serve as the primary validated cases, while the remaining languages provide exploratory cross-script and cross-family comparisons. Across this corpus, Bengali requires (1.56\times) as many GPT-4o tokens as English, reducing a nominal 128k-token context window to an effective 82k-token English-equivalent capacity for the same semantic content. With the Qwen2.5 and Mistral tokenizers, Bengali requires up to (4.5\times) the English token count. Yoruba, despite using the Latin script, exhibits the highest GPT-4o tokenization premium at (2.37\times), indicating that tokenization inequity cannot be explained by script family alone. These results demonstrate that tokenization can create measurable economic and functional barriers, highlighting the need to treat tokenization as an equity-relevant infrastructure layer for underserved language communities, particularly where educational systems depend on low-cost or offline-capable AI tools.

CommentsAccepted at IJCAI 2026 Workshop (https://lm4uc.github.io/)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑