无形的语言税:2026年大语言模型分词器中法语及地区语言的令牌溢价,以及一个法语优化原型
The Invisible Language Tax: Token Premiums of French and Regional Languages in 2026 LLM Tokenizers, and a French-Optimized Prototype
浏览论文内容
中文总结 AI 辅助
本研究测量了2026年主流大语言模型分词器中法语的令牌溢价(比英语多31%-58%),并提出了法语优化的字节级BPE原型Baracoda FR v1.2,在法语上比Tekken少用11.5%的令牌,但其他语言表现较差。
中文摘要 AI 辅助
大语言模型服务按令牌计费,上下文窗口以令牌为单位衡量,然而同一内容在不同语言中所需的令牌数量却存在差异。我们在七个广泛使用的2026年模型分词器(OpenAI o200k、Llama 3、Qwen3、DeepSeek V3/V4、Gemma 3、Mistral Tekken,以及通过Anthropic计数API使用的Claude第5代分词器)上,针对NTREX-128(124个非英语参考译文)和《世界人权宣言》的地区语言版本,测量了这种令牌溢价。法语比英语多需要31%至58%的令牌,而简体中文则从少5%到多40%不等,且在七个分词器中的六个上比法语更便宜。法国的地区语言和海外语言支付的令牌数量约为英语的1.6至3.3倍。我们讨论了历史重发、分层定价和固定上下文窗口如何放大智能体使用中的绝对差距。在一项受控实验(BPE、Europarl、50k词汇量)中,将法语加入分词器训练数据可迅速降低溢价,但收益递减,且英语的成本随之增加。最后,我们提出了Baracoda FR v1.2,一个采用Tekken词汇量大小的字节级BPE原型。在对六个设计期间从未参考过的语料库进行的最终测试中,采用事先固定的协议,该原型在法语上比Tekken少使用11.5%的令牌,在英语上少3.7%;在移除与训练数据重叠的测试句子后,以及在使用相等的普通令牌预算时,结果依然成立。它在其他语言上表现较差,且在相当的词汇量下,并未超越CroissantLLM。这些仅是分词结果;对模型质量和任务成本的影响仍有待证明。
英文摘要
LLM services are billed per token and context windows are measured in tokens, yet the number of tokens needed for the same content varies across languages. We measure this token premium on seven tokenizers of widely used 2026 models (OpenAI o200k, Llama 3, Qwen3, DeepSeek V3/V4, Gemma 3, Mistral Tekken, and the Claude generation-5 tokenizer via Anthropic's counting API) on NTREX-128 (124 non-English reference translations) and on the Universal Declaration of Human Rights for regional languages. French requires 31% to 58% more tokens than English, whereas Simplified Chinese ranges from 5% fewer to 40% more and is cheaper than French on six of the seven tokenizers. Regional and overseas languages of France pay roughly 1.6 to 3.3 times the English count. We discuss how history re-sending, tiered pricing and fixed context windows amplify the absolute gap in agentic use. In a controlled experiment (BPE, Europarl, 50k vocabulary), adding French to tokenizer training data quickly reduces the premium, with diminishing returns and a growing cost for English. Finally, we present Baracoda FR v1.2, a byte-level BPE prototype with Tekken's vocabulary size. On a final test of six corpora never consulted during design, with a protocol declared fixed beforehand, it uses 11.5% fewer tokens than Tekken on French and 3.7% fewer on English; results hold after removing test sentences overlapping the training data and with an equal ordinary-token budget. It is worse on other languages and, at comparable vocabulary size, does not outperform CroissantLLM. These are segmentation results only; effects on model quality and task cost remain to be shown.
发表机构
- Baracoda AI Labs(Baracoda AI实验室)
- NEOMA Business School(NEOMA商学院)
机构由 AI 辅助整理,请以论文原文为准。