arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00378cs.CL

可移除且不可约:多语言分词税的代币成本账本

Removable and Irreducible: A Token-Cost Ledger for the Multilingual Tokenization Tax

Madhulatha Mandarapu, Sandeep Kunkunuru

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大型语言模型的多语言分词税,构建代币成本账本拆分成本来源,通过实验发现可移除大部分超额代币成本,为相关研究提供统一框架与开源工具。

中文摘要 AI 辅助

大型语言模型对非英语文本存在公认的“税”:相同内容的代币成本是英语的数倍,且由于注意力机制的计算量与序列长度呈二次关系,所需计算资源也大幅增加。我们探究该“税”有多少是可移除的。将代币层视为源编码——Transformer的计算量随序列长度单调变化,其每个原子的下限为香农率$H/\log_2 V$,这一概念已在先前研究中应用于分词器——我们构建了一个代币成本账本,在固定平行内容下,将每种语言的成本拆分为可移除的编码冗余、剩余编码 slack、固有内容项,以及正交且不可约的音素- grapheme转换项,后者控制的是多模态而非文本成本。在FLORES-200数据集的8种语言上,某生产级分词器对印度文字的代币成本最高是英语的8.9倍;基于1012个句子训练的匹配脚本代码可移除该超额成本的中位数64%(自助法95%置信区间[0.638, 0.647]),且脚本公平信息下限显示固有内容差异不足6%——该“税”属于表征层面,而非信息层面。构建的代码可移除受控源冗余的98%,且代币“税”意味着注意力成本最高可达79倍。我们明确研究范围与局限性:这是计算与内存的核算,而非模型质量的断言;我们未测量也未断言正字法项的跨语言方向;且匹配代码是保守的小数据演示。我们提供统一账本、可移除与固有归因方法,以及开源的单命令工具。

英文摘要

Large language models pay a well-documented tax on non-English text: the same content costs several times more tokens, and because attention is quadratic in sequence length, far more compute. We ask how much of this tax is removable. Framing the token layer as source coding -- transformer compute is monotone in sequence length, whose per-atom floor is the Shannon rate $H/\log_2 V$, an object already applied to tokenizers in prior work -- we assemble a token-cost ledger that splits each language's cost, at fixed parallel content, into a removable coding redundancy, a residual coding slack, an intrinsic-content term, and an orthogonal, irreducible grapheme-to-phoneme term that governs the multimodal rather than the text cost. On FLORES-200 across eight languages, a production tokenizer costs up to $8.9\times$ more tokens for Indic scripts than for English; a script-matched code trained on $1,012$ sentences removes a median $64\%$ of that excess (bootstrap 95\% CI $[0.638, 0.647]$), and a script-fair information floor shows the intrinsic content differs by under $6\%$ -- the tax is representational, not informational. A constructed code removes $98\%$ of a controlled source's redundancy, and the token tax implies up to $79\times$ attention cost. We are explicit about scope and failure: this is compute-and-memory accounting, not a model-quality claim; we neither measure nor claim the cross-lingual direction of the orthographic term; and our matched code is a conservative small-data demonstration. We contribute the unifying ledger, the removable-versus-intrinsic attribution, and an open one-command harness.

发表机构

  • VaidhyaMegha Private Limited(VaidhyaMegha私人有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑