保留语素:面向尼泊尔语的形态学引导预分词
Preserving Morphemes: Morphology-Guided Pre-Tokenization for Nepali
查看机构详情
- Kathmandu University(加德满都大学)
- IOE Thapathali Campus(IOE塔帕塔利校区)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出形态学引导的预分词器Papaya,在BPE前拆分尼泊尔语单词为词干和词缀,以保留语素完整性,在分词和语言建模上略有提升,但下游任务影响有限。
中文摘要 AI 辅助
字节级BPE词汇表将尼泊尔语单词的每种屈折形式学习为单独的字符串,因此名词词干在其各种带格标记的形式中拼写不同。我们测试了在BPE之前将单词拆分为词干和词缀是否有帮助,同时保持语料库、词汇表大小、模型和训练步数不变。我们的预分词器Papaya使用基于已发表的尼泊尔语语法构建的有限状态转换器,回退到正则表达式,并保持BPE训练器不变。在由七位母语者标注的607个单词上,其分词器达到0.96的边界F1分数,且生成的标记比普通BPE更频繁地保持词干完整。在17M参数的语言模型中,在相同训练步数下,每字节比特数降低了约1%;在相同轮数下观察到的较大增益大部分来自较长标记序列带来的额外步数,而无监督的Morfessor分词也给出了相同的改进。在下游任务中效果较小:NER仅在包含训练中未见单词的实体上有所提升,词性标注和新闻分类没有变化,已发布的尼泊尔语分词器表现相当。我们发布了带标注的边界集、包含556个词缀的数据集以及代码。
英文摘要
A byte-level BPE vocabulary learns each inflected form of a Nepali word as a separate string, so a noun stem is spelled differently in each of its case-marked forms. We test whether splitting words into stem and affixes before BPE helps, with the corpus, vocabulary size, model and number of training steps held fixed. Our pre-tokenizer, Papaya, uses a finite-state transducer built from a published grammar of Nepali, falls back to regular expressions, and leaves the BPE trainer unchanged. On 607 words annotated by seven native speakers its segmenter reaches 0.96 boundary F1, and the resulting tokens keep stems intact far more often than plain BPE does. In a 17M-parameter language model it lowers bits per byte by about 1% at equal training steps; most of the larger gain seen at equal epochs comes from the extra steps that longer token sequences buy, and an unsupervised Morfessor segmentation gives the same improvement. Downstream the effect is small: NER improves only on entities that contain words unseen in training, POS tagging and news classification do not change, and published Nepali tokenizers perform about as well. We release the annotated boundary set, a 556-affix dataset and the code.