arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当分词与语言建模联合优化时,会学习到哪些词元?

What Tokens are Learned when Tokenization is Optimized Jointly with Language Modeling?

Saketh Reddy Vemula, Parameswari Krishnamurthy

arXiv 2608.17325首次发表:更新:

发表机构

IIIT Hyderabad(印度国际信息技术研究院(海得拉巴))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究对比无分词器方法(SSLMs、H-Nets)与固定分词器在18种语言上的效果,发现联合优化分词与语言建模可生成高效且有效的差异化词汇,为下游NLP提供新方案。

AI 中文摘要

分词是语言建模流程的基础组件,尽管其重要性不言而喻,但通常被固定,即便它会显著影响不同语言的模型性能。本研究分析了分词与语言建模联合优化时会学习到哪些词元,在18种类型学和书写系统各异的语言上,将无分词器方法(如SSLMs和H-Nets)与固定分词器进行了比较。结果表明,联合优化从根本上改变了词元结构:SSLMs恢复了形态对齐且上下文高效的词元,而H-Nets优先考虑字节级效率,生成的词元更长,与标准子词词汇的重叠度极低。我们还发现,分词行为因语言类型而异:黏着语在学习过程中表现出更动态的分词模式。通过下游评估,采用预训练后微调的BERT模型,我们发现基于SSLM的预分词始终能降低语言建模困惑度,且尽管词汇不同,仍能取得具有竞争力的下游性能。总体而言,无分词器方法优化的是上下文和计算效率,而非严格的形态结构,从而为下游NLP生成了截然不同但有效的词汇表。

英文摘要

Tokenization is a fundamental component of language modeling pipelines. Despite its importance, it is often fixed, even though it significantly impacts model performance across languages. In this work, we analyze what tokens are learned when tokenization is jointly optimized with language modeling. We compare tokenizer-free approaches such as SSLMs and H-Nets with fixed tokenizers across 18 typologically and script-diverse languages. Our results show that joint optimization fundamentally alters token structure. SSLMs recover morphologically aligned and contextually efficient tokens, whereas H-Nets prioritize byte-level efficiency, producing longer tokens with very low overlap with standard subword vocabularies. We further show that tokenization behavior varies across language typologies. Agglutinative languages exhibit more dynamic segmentation patterns while learning. Through downstream evaluation, with pretrained-then-finetuned BERT models, we find that SSLM-based pretokenization consistently reduces language modeling perplexity and achieves competitive downstream performance despite distinct vocabularies. Overall, tokenizer-free approaches optimize for contextual and computational efficiency rather than strict morphological structure, resulting in fundamentally different yet effective vocabularies for downstream NLP.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑