发表机构
Microsoft Research Africa; Microsoft Research Accelerator; Microsoft Research India(微软研究院非洲分部; 微软研究院加速器; 微软研究院印度分部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出与语言无关的潜在核心分词器(LCT),通过分离结构发现与词汇构建,在104种语言的分词任务中优于BPE等方法,多语言下游基准得分提升,证明压缩 alone 无法预测表示质量。
AI 中文摘要
分词器通常以压缩为优化目标,但紧凑的词汇表并不一定能在各语言间均匀分配容量。我们提出潜在核心分词器(Latent Core Tokenizer,LCT),一种与语言无关的方法,将结构发现与词汇构建分离。LCT采用最小描述长度、基于熵的边界信号和形态句法约束,在构建共享词汇表前识别可复用的语言单元。在104种语言、20万词元词汇表的设置下,LCT的词元丰度更低、形态分数(MorphScore)更高,优于BPE、Unigram和感知奇偶性的BPE,同时保持分词成本的跨语言差异相当。在四个多语言下游基准任务中,LCT的综合得分较BPE、Unigram和感知奇偶性的BPE分别提升1.48、1.83和2.00分。我们的研究表明,仅靠压缩无法预测表示质量,凸显了形态驱动的结构发现以及频率在各语言间分配最终词汇表的重要性。
英文摘要
Tokenizers are commonly optimized for compression, but a compact vocabulary does not necessarily distribute its capacity evenly across languages. We introduce the Latent Core Tokenizer (LCT), a language-agnostic approach that separates structural discovery from vocabulary construction. LCT uses Minimum Description Length, entropy-based boundary signals, and morphotactic constraints to identify reusable linguistic units before constructing a shared vocabulary. Across 104 languages with a 200K-token vocabulary, LCT achieves lower fertility and higher MorphScore than BPE, Unigram, and parity-aware BPE, while maintaining comparable cross-lingual disparity in tokenization cost. Across four multilingual downstream benchmarks, LCT improves aggregate score by 1.48, 1.83, and 2.00 points over BPE, Unigram, and parity-aware BPE, respectively. Our findings show that compression alone does not predict representation quality and highlight the importance of morphology-driven structural discovery and how frequency is used to allocate the final vocabulary across languages.
CommentsUnder review