婆罗米系文字的类型驱动分词
Type-Driven Tokenization for Brahmic Scripts
浏览论文内容
中文总结 AI 辅助
针对大型语言模型在婆罗米系文字上分词产生格式错误的问题,本文形式化区分了英语半群与婆罗米偏半群的正字法结构,推导出可证明正确的fixToken函数,并转化为SentencePiece补丁及Rust预分词器库,消除了印度诸文字的分词错误。
中文摘要 AI 辅助
大型语言模型中使用的标准分词器在应用于婆罗米系文字时会产生格式错误的文本。婆罗米系文字是一类元音附标文字(abugidas),其书写系统中辅音带有内在元音,且该元音可被依附标记修改。这类文字包括天城文(Devanagari)、泰卢固文(Telugu)、泰米尔文(Tamil)、卡纳达文(Kannada)等。根本问题在于,这些分词器违反了在英语等字母文字中不存在的正字法约束。我们观察到,英语正字法构成一个“半群”(任意两个有效词元可自由拼接),而婆罗米系文字正字法构成一个“偏半群”:并非所有拼接都能产生有效字符串。我们在Agda中形式化了这一区别,将有效的婆罗米系词元建模为转换系统中的链,并推导出一个可证明正确的fixToken函数,该函数扩展任何候选词元以尊重正字法边界。随后,我们展示了这一形式化推导如何转化为对SentencePiece的实用补丁,以及一个独立的基于Rust的预分词器库,从而消除了印度诸文字中观察到的错误。
英文摘要
Standard tokenizers used in large language models produce malformed text when applied to Brahmic scripts. They are a family of abugidas, writing systems whose consonants carry an inherent vowel that dependent marks can modify. They include Devanagari, Telugu, Tamil, Kannada, and others. The underlying issue is that these tokenizers violate orthographic constraints that do not arise in alphabetic scripts like English. We observe that while English orthography forms a \emph{semigroup} (any two valid tokens can be freely concatenated), Brahmic orthography forms a \emph{partial semigroup}: not every concatenation yields a valid string. We formalise this distinction in Agda, model valid Brahmic tokens as chains in a transition system, and derive a provably correct \texttt{fixToken} function that extends any candidate token to respect orthographic boundaries. We then show how this formal derivation translates into a practical patch for SentencePiece as well as a standalone Rust-based pre-tokenizer library, eliminating the observed errors across Indic scripts.
发表机构
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。