arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨语言同形异义词的分词

Tokenizing Crosslingual Homographs

Rotem Brillant, Yuval Pinter

arXiv 2607.17689首次发表:更新:

AI 中文总结

研究跨语言同形异义词在多语言模型中的处理问题,提出基于语言线索的分词器干预方法,通过内在分析和机器翻译实验验证,该方法能消除不同分词器处理差异,在部分设置下有改进,为添加语言信息提供探索方向。

AI 中文摘要

多语言语言模型依靠共享子词词汇表在有限数量的词元单元中表示多种语言。虽然这种共享通常很有用,但也会导致相同表面形式在不同语言中被过度统一处理,即便其含义或用法不同。我们通过跨语言同形异义词和伪朋友来研究这一局限性,并探讨在分词过程中更早引入语言信息是否能改善这种情况。我们提出了一种基于语言线索的简单分词器层面的干预方法:用特定语言字符替换共享词汇单词的首字符,在词汇构建过程中减少共性。在内在分析中,我们通过分词器层面的统计发现,BPE和UnigramLM通常以基本不考虑语言的方式处理跨语言同形异义词,而上下文敏感的SaGe分词器差异更大;我们的干预消除了这种差距。在下游英语到X的机器翻译中,我们的线索在几种情况下有适度改进,特别是在BPE下,尽管效果在所有语言和评估集上并不一致。总体而言,研究结果表明在分词器层面添加轻量级语言信息是一个有前景的进一步探索方向。

英文摘要

Multilingual language models rely on shared subword vocabularies to represent multiple languages within a limited number of token units. While such sharing is often useful, it can also create cases in which identical surface forms are treated too uniformly across languages, even when their meanings or usage differ. We investigate this limitation through cross-lingual homographs and false friends, and examine whether introducing language information earlier in the tokenization process can improve their treatment. We propose a simple tokenizer-level intervention based on language cues: language-specific characters replacing initial characters of shared-vocabulary words, reducing common identity during vocabulary construction. In intrinsic analysis, we find through tokenizer-level statistics that BPE and UnigramLM often treat cross-lingual homographs in a largely language agnostic way, whereas the context-sensitive SaGe tokenizer diverges more strongly; our intervention removes this gap. In downstream English-to-X machine translation, our cues yield modest improvements in several settings, especially under BPE, although the effect is not consistent across all languages and evaluation sets. Overall, the findings suggest that adding lightweight language information at the tokenizer level is a promising direction for further exploration.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑