泰米尔语的字素感知印度语分词器:大规模训练与内在评估
A Grapheme-Aware Indic Tokenizer for Tamil: Large-Scale Training and Intrinsic Evaluation
浏览论文内容
中文总结 AI 辅助
本文提出一种针对泰米尔语的字素感知分词器,通过可逆Unicode映射保留完整字素簇,在压缩比和碎片化等内在指标上优于五种主流多语言分词器。
中文摘要 AI 辅助
分词是现代自然语言处理(NLP)系统的基础,它将原始文本转换为神经语言模型可以处理的离散单元。这一过程的有效性直接影响词汇效率、序列长度、计算成本和下游模型性能。尽管诸如字节对编码(BPE)、WordPiece和SentencePiece等多语言分词器在众多语言中表现良好,但它们对形态丰富的印度语系语言的分割往往效率低下。泰米尔语尤其面临独特挑战,因为其基于字素的书写系统可以用多个Unicode码点表示一个可见字符。在本工作中,我们提出了一种针对泰米尔语的字素感知印度语分词器,该分词器在WordPiece词汇学习之前,通过可逆的Unicode映射策略保留完整的字素簇。通过操作字素级表示而非单个Unicode码点,该分词器产生了具有语言学意义的词元边界,同时保持与基于Transformer的语言模型的完全兼容性。该分词器在大规模泰米尔语语料库上进行训练,并使用一个全面的内在评估框架进行评估,该框架衡量压缩效率、词元碎片化、信息密度和词汇利用率。实验评估将所提出的分词器与五种广泛使用的多语言分词器进行比较:GPT-2、mBERT、mT5、mBART和NLLB。所提出的分词器在面向碎片化和序列效率的内在指标上,在评估的分词器中取得了最强性能,同时匹配了观察到的最高压缩比。这些结果证明了字素感知预处理对泰米尔语分词的有效性。
英文摘要
Tokenization forms the foundation of modern Natural Language Processing (NLP) systems by transforming raw text into discrete units that neural language models can process. The effectiveness of this process directly influences vocabulary efficiency, sequence length, computational cost, and downstream model performance. Although multilingual tokenizers such as Byte Pair Encoding (BPE), WordPiece, and SentencePiece have performed well across numerous languages, they often segment morphologically rich Indic languages inefficiently. Tamil, in particular, poses unique challenges because its grapheme-based writing system can represent a single visible character with multiple Unicode code points. In this work, we present a grapheme-aware Indic tokenizer for Tamil that preserves complete grapheme clusters through a reversible Unicode mapping strategy prior to WordPiece vocabulary learning. By operating on grapheme-level representations instead of individual Unicode code points, the tokenizer produces linguistically meaningful token boundaries while remaining fully compatible with transformer-based language models. The tokenizer is trained on a large-scale Tamil corpus and evaluated using a comprehensive intrinsic evaluation framework that measures compression efficiency, token fragmentation, information density, and vocabulary utilization. Experimental evaluation compares the proposed tokenizer against five widely used multilingual tokenizers: GPT-2, mBERT, mT5, mBART, and NLLB. The proposed tokenizer achieves the strongest performance among the evaluated tokenizers on fragmentation- and sequence-efficiency-oriented intrinsic metrics, while matching the highest observed compression ratio. These results demonstrate the effectiveness of grapheme-aware preprocessing for Tamil tokenization.
发表机构
- IIITDM Kanchipuram(印度信息技术与设计学院坎奇普兰分校)
- Wadhwani School of Data Science and AI, IIT Madras(印度理工学院马德拉斯分校瓦德瓦尼数据科学与人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。