arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

词汇规范化中的多语言诅咒

The Curse of Multilinguality in Lexical Normalization

Saman Rahbar

arXiv 2609.00329首次发表:更新:

AI 中文总结

本文研究词汇规范化中多语言模型的最优训练语言数量,发现仅用1-4种语言联合训练的字符级模型准确率最高,多语言竞争导致准确率下降,紧凑模型需少语言训练。

AI 中文摘要

词汇规范化是将用户生成文本中充满噪声的非标准词汇(如tmrw、u、gr8)改写为标准形式的任务。由于大多数语言的标注数据稀缺,一种流行的捷径是在多种语言上同时训练单个模型。本文提出一个简单问题:此类模型应在多少种语言上进行训练?本文采用固定容量的字符级模型和标准基准中的12种语言,将联合训练的语言数量从1种到12种变化,并测量每种语言的准确率。研究发现存在明显的多语言诅咒现象:当一种语言仅与少数其他语言(通常为1到4种)一起训练时,准确率最高;随着加入的语言增多,准确率会稳步且大幅下降,当加入其余语言时,准确率下降约40%。控制总训练数据量不变的对照实验显示,准确率下降出现得更早且幅度更大,这表明问题源于多种语言在固定大小的模型中存在竞争,而非可用数据量的问题。本文还测试一种语言与其他语言的类型学距离是否可预测其理想的联合训练语言数量,未发现可靠规律:任何明显的关系都依赖于少数几种语言,无法成立。对于紧凑的规范化模型,少即是多:少数语言的训练效果优于将所有语言合并到单个模型中。

英文摘要

Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u, gr8) into their standard forms. Because labelled data is scarce for most languages, a popular shortcut is to train a single model on many languages at once. We ask a simple question: how many languages should such a model be trained on? Using one fixed-capacity character-level model and twelve languages from a standard benchmark, we vary the number of jointly trained languages from one to twelve and measure per-language accuracy. We find a clear curse of multilinguality: accuracy is highest when a language is trained with only a few others, often just one to four, and then falls steadily and substantially, dropping by about forty percent as the rest are piled on. A control that holds the total amount of training data constant makes the decline arrive sooner and fall further, which points to competition among the languages for one fixed-size model rather than to how much data is available. We also test whether a language's typological distance from the others predicts its ideal number of co-training languages, and find no dependable rule: any apparent relationship rests on a couple of languages and does not hold up. For compact normalization models, less can be more: a few languages beat pooling everything into a single model.

CommentsAccepted at the 11th Workshop on Natural User-generated Text (W-NUT 2026), EMNLP 2026. Camera-ready version. 7 pages, 3 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑