发表机构
Centre for Sensors, Instrumentation and Cyber-Physical System Engineering (Centre for SeNSE), Indian Institute of Technology Delhi; RSL Quantum, FITT, IIT Delhi(印度理工学院德里分校传感器、仪器仪表和网络物理系统工程中心(SeNSE中心); 印度理工学院德里分校FITT的RSL量子公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对古典印度语言分词低效问题,提出BHARATI,它基于多语言语料库训练出三个版本的分词器,经子词丰富度分析和实验验证,v3表现出色,能减少序列长度,增加下游语言模型有效上下文长度,相关模型、脚本和基准已开源。
AI 中文摘要
标准的子词分词算法,如字节对编码(BPE)和SentencePiece,主要在现代语言语料库上训练,应用于古典印度语言时会产生低效的分词结果。梵语、泰米尔语和其他古典印度语言具有黏着形态、丰富的连声现象(词边界处的音位融合)以及通用训练数据中没有的特定领域词汇。本文提出了BHARATI,这是一组基于781MB平衡语料库训练的SentencePiece BPE分词器,该语料库涵盖七种语言(英语、印地语、梵语、泰米尔语、泰卢固语、卡纳达语和马拉雅拉姆语),并支持所有语言的母语脚本。我们描述了三个连续的分词器版本:v1(仅英语和梵语,泰米尔语采用字节回退)、v2(四种语言支持,南方语言采用字节级回退)和v3(完整的七种语言母语子词覆盖)。子词丰富度分析表明,v3对每个印度知识系统(IKS)技术术语平均产生2.6个词元,而GPT - 2的分词器每个术语平均产生5.25个词元,多语言SentencePiece基线为3.75个词元,在一组保留的IKS术语上收益最大,这些术语按构造表示为单个词元。在490个IKS领域句子的保留测试集上(七种语言每种70个句子,随测量脚本发布),v3相对于GPT - 2和字节级编码(缺乏印度语母语子词)将序列长度减少了约90%,相对于mBART - 50多语言基线减少了约25%,六种印度语言平均计算,这直接转化为下游语言模型有效上下文长度的增加。分词器模型(32000词汇量)、训练脚本和评估基准已在开放许可下发布。
英文摘要
Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages. Sanskrit, Tamil, and other classical Indic languages exhibit agglutinative morphology, productive sandhi (phonological fusion at word boundaries), and domain-specific vocabularies absent from general-purpose training data. This paper presents BHARATI, a set of SentencePiece BPE tokenizers trained on a balanced 781 MB corpus spanning seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam) with native script support for all languages. We describe three successive tokenizer versions: v1 (English and Sanskrit only, with broken byte-fallback for Tamil), v2 (four-language support with byte-level fallback for southern languages), and v3 (full seven-language native subword coverage). Subword fertility analysis demonstrates that v3 averages 2.6 tokens per Indian Knowledge System (IKS) technical term, compared to 5.25 tokens per term with GPT-2's tokenizer and 3.75 tokens with the multilingual SentencePiece baseline, with the largest gains on a set of reserved IKS terms that are represented as single tokens by construction. On a held-out test set of 490 IKS-domain sentences (70 per language across seven languages, released with the measurement script), v3 reduces sequence length by roughly 90% relative to GPT-2 and byte-level encoding (which lack native Indic subwords) and by approximately 25% relative to the mBART-50 multilingual baseline, averaged across the six Indic languages, directly translating to increased effective context length for downstream language models. The tokenizer models (32,000 vocabulary), training scripts, and evaluation benchmarks are released under open licenses.
Comments33 pages, 6 figures