arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向泰米尔语语言模型的形态学感知可逆语义分词与分层词组合

Morphology Aware Reversible Semantic Tokenization and Hierarchical Word Composition for Tamil Language Models

Anand Murugan

arXiv 2608.01153首次发表:更新:

AI 中文总结

该研究针对泰米尔语,提出结合ThamizhiMorph的形态系统与分层词组合方法,在固定小模型预算下提升翻译性能,同时大幅降低序列长度与推理成本。

AI 中文摘要

统计子词分词器可处理任意文本,但其单元未必与词汇或语法结构对齐,这对泰米尔语尤为重要——泰米尔语的书面词可编码词干变化、格、数、时态、一致关系、语态、附着词及关联动词。我们提出一套泰米尔语形态系统,扩展了开源ThamizhiMorph分析器与生成器,同时结合字节精确语义分词器及学习型分层词组合器。12个有限状态转换器将词分析为词干与语法特征,字符与字节 fallback机制确保精确重构。我们对比了扁平形态分词器、信号保留型词组合器,以及基于Sarvam-1、AI4Bharat IndicBERTv2、BrahmicTokenizer-131K的分词器,所有系统均使用相同的69591个泰米尔语-英语训练对、1897万参数的编码器-解码器、40000次更新、目标分词器、优化器、位置方法及生成设置。在受保护的3539行IN22与FLORES+评估中,形态扁平分词器取得最佳综合得分:BLEU为10.63,chrF++为35.26,COMETKiwi为0.6276;相较于最强外部分词器基线AI4Bharat,分别提升7.2%、3.2%、2.6%。词组合器得分分别为10.30、34.88、0.6241,较AI4Bharat分别提升3.8%、2.1%、2.0%;其将平均全局源状态从71.48降至29.08,降幅59.3%,且根据解码器缓存情况,估计可减少9%-21%的推理FLOPs。其剩余质量差距集中在较长的FLORES+句子中。这些结果表明,在固定小模型预算下,显式泰米尔语形态可提升翻译性能,而分层组合可大幅减少序列长度与估计推理成本。

英文摘要

Statistical subword tokenizers can process arbitrary text, but their units need not align with lexical or grammatical structure. This is especially important for Tamil, where a written word may encode stem changes, case, number, tense, agreement, voice, clitics, and linked verbs. We present a Tamil morphology system extending the open-source ThamizhiMorph analyzer and generator, together with a byte-exact semantic tokenizer and a learned hierarchical word composer. Twelve finite-state transducers analyze words into lemmas and grammatical features, while character and byte fallbacks preserve exact reconstruction. We compare a flat morphology tokenizer, a signal-preserving word composer, and tokenizers based on Sarvam-1, AI4Bharat IndicBERTv2, and BrahmicTokenizer-131K. All systems use the same 69,591 Tamil-English training pairs, 18.97-million-parameter encoder-decoder, 40,000 updates, target tokenizer, optimizer, positional method, and generation settings. On a protected 3,539-row IN22 and FLORES+ evaluation, morphology-flat achieves the best pooled scores: 10.63 BLEU, 35.26 chrF++, and 0.6276 COMETKiwi. Relative to AI4Bharat, the strongest external-tokenizer baseline, these are improvements of 7.2%, 3.2%, and 2.6%. The word composer scores 10.30, 34.88, and 0.6241, improving on AI4Bharat by 3.8%, 2.1%, and 2.0%. The composer reduces mean global source states from 71.48 to 29.08, a 59.3% reduction, and is estimated to require 9-21% fewer inference FLOPs depending on decoder caching. Its remaining quality gap is concentrated in longer FLORES+ sentences. These results show that explicit Tamil morphology improves translation under a fixed small-model budget, while hierarchical composition substantially reduces sequence length and estimated inference cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑