arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2508.06533cs.CLcs.AI

断词的艺术:重新思考多语言分词器设计

The Art of Breaking Words: Rethinking Multilingual Tokenizer Design

  • BharatGen Team(BharatGen团队)

机构由 AI 辅助整理,请以论文原文为准。

Aamod Thakur, Ajay Nagpal, Atharva Savarkar, Kundeshwar Pundalik, Siddhesh Dosi, Piyush Sawarkar, Viraj Thakur, Rohit Saluja, Maunendra Sankar Desarkar, Ganesh Ramakrishnan

更新

AI总结:

本文系统研究了多语言分词器设计中词汇表大小、预分词规则和训练语料组成对效率与模型质量的影响,提出一种平衡多语言数据的组成算法,在印度系文字上显著降低了词元与单词比率并提升了推理速度。

AI中文摘要:

尽管模型架构和训练目标已被充分研究,但分词,特别是在多语言环境下,仍是大型语言模型(LLM)开发中相对被忽视的方面。现有的分词器通常表现出高词元与单词比率、上下文长度的低效使用以及较慢的推理速度。我们提出了一项系统研究,将词汇表大小、预分词规则和训练语料组成与词元与单词效率及模型质量联系起来。为了将我们的分析置于语言多样化的语境中,我们对印度系文字进行了广泛的实验,这些文字由于其高度的字体多样性和正字法复杂性而呈现出独特的挑战。基于这些分析的见解,我们提出了一种新颖的数据组成算法,该算法平衡了用于分词器训练的多语言数据。我们对预分词策略的观察显著提高了模型性能,并且我们的数据组成算法相对于传统的数据随机化方法将平均词元与单词比率降低了约6%。我们的分词器在平均词元与单词比率上比最先进的多语言印度系模型实现了超过40%的改进。这种改进在模型性能和推理速度上产生了可衡量的收益。这突显了分词与架构和训练目标一起,是构建高效、可扩展多语言LLM的关键杠杆。

英文摘要:

While model architecture and training objectives are well-studied, tokenization, particularly in multilingual contexts, remains a relatively neglected aspect of Large Language Model (LLM) development. Existing tokenizers often exhibit high token-to-word ratios, inefficient use of context length, and slower inference. We present a systematic study that links vocabulary size, pre-tokenization rules, and training-corpus composition to both token-to-word efficiency and model quality. To ground our analysis in a linguistically diverse context, we conduct extensive experiments on Indic scripts, which present unique challenges due to their high script diversity and orthographic complexity. Drawing on the insights from these analyses, we propose a novel algorithm for data composition that balances multilingual data for tokenizer training. Our observations on pretokenization strategies significantly improve model performance, and our data composition algorithm reduces the average token-to-word ratio by approximately 6% with respect to the conventional data randomization approach. Our tokenizer achieves more than 40% improvement on average token-to-word ratio against stateof-the-art multilingual Indic models. This improvement yields measurable gains in both model performance and inference speed. This highlights tokenization alongside architecture and training objectives as a critical lever for building efficient, scalable multilingual LLMs

↑