发表机构
IIT Bombay(印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究发现语言模型对非规范分词的鲁棒性因语言而异,Llama-3.1-8B等模型在高碎片化语言中性能下降,LoRA微调可有效缓解分词敏感性。
AI 中文摘要
给定字符串存在指数级多的有效分词方式,但语言模型仅使用分词器确定性生成的单一规范序列,对更广泛的分词空间缺乏研究。本文针对这一被忽视的空间,研究语言模型在跨语言非规范分词下的行为。此前研究表明,英文模型对代表同一底层字符串的替代分词基本不变,本文探究这种不变性是否适用于非英语语言。我们对涵盖不同文字的27种语言开展多语言研究,在6项下游任务中评估大语言模型(LLM)在替代分词下的表现。结果显示,分词不变性不具普适性:不同语言的模型行为差异显著,指令微调模型中,Llama-3.1-8B的平均相对性能下降23.7%,Qwen3-8B为11.4%,Gemma-3-12B为9.9%。分词不变性的变化在语言间具有系统性:分词碎片化程度更高的语言对非规范分词的敏感性显著更强。本文对分词鲁棒性的研究可用于诊断模型与分词器的耦合紧密程度,证明分词鲁棒性并非语言模型的普遍属性,而是高度依赖语言及其与分词器的交互。我们还发现,采用多分词训练数据的LoRA微调可有效缓解分词敏感性:仅在英文上微调可提升跨语言分词鲁棒性,而系统采样多样化非规范分词可实现最强整体性能。
英文摘要
Despite the existence of exponentially many valid tokenizations for a given string, language models operate on a single canonical sequence deterministically produced by the tokenizer, leaving the broader tokenization space largely uncharacterized. In this paper, we investigate this overlooked space by studying the behavior of language models under non-canonical tokenizations across diverse languages. For English, prior work shows that models are largely invariant to alternative tokenizations that represent the same underlying string. We ask whether this invariance generalizes to other languages beyond English. We conduct a multilingual study across 27 languages spanning diverse scripts and evaluate LLM behavior under alternative tokenizations across six downstream tasks. We find that tokenization invariance does not generalize: model behavior varies substantially across languages with instruction-tuned models exhibiting an average relative performance drop of 23.7% for Llama-3.1-8B, 11.4% for Qwen3-8B, and 9.9% for Gemma-3-12B. The variation of tokenization invariance is systematic across languages. Languages that exhibit higher token fragmentation show significantly greater sensitivity to non-canonical tokenizations. Our study of tokenization robustness serves as a diagnostic of how tightly a model is coupled to its tokenizer. These results demonstrate that tokenization robustness is not a universal property of language models, but depends strongly on the language and its interaction with the tokenizer. We also show that LoRA fine-tuning with multi-tokenization training data provides an effective mitigation for tokenization sensitivity. Fine-tuning on English alone improves tokenization robustness across languages, while systematically sampling diverse non-canonical tokenizations achieves the strongest overall performance.