MoirfEolas与CríochScore:开发面向爱尔兰语的资源及与爱尔兰语形态学的分词对齐评估
MoirfEolas and CríochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology
浏览论文内容
中文总结 AI 辅助
本文开发了爱尔兰语形态学分词资源MoirfEolas和评估指标CríochScore,经评估发现Unigram Language Model分词算法与爱尔兰语形态对齐性更优,相关成果助力爱尔兰语NLP发展且可推广至其他语言。
中文摘要 AI 辅助
本文提出了面向爱尔兰语的新型分词资源以及与该语言形态边界对齐的评估指标。我们推出了MoirfEolas,这是一个包含超过35000个爱尔兰语单词的数据集,这些单词被映射到各自的元音省略(eclipses)、前缀和后缀;同时提出了评估指标CríochScore,用于评估分词与MoirfEolas中存在的形态边界的对齐程度。我们使用CríochScore以及分词文献中现有的内在指标对常见的分词算法进行评估,发现Unigram Language Model(一元语言模型)比其他被评估的算法更常与爱尔兰语形态学对齐。我们还发现分词的形态对齐与压缩率、词汇效率之间存在权衡,为爱尔兰语自然语言处理开发提供了实用见解。该数据集有助于改善爱尔兰语的低资源状态;此外,本文报告的构建过程可被其他语言效仿,以创建专门的形态学资源。
英文摘要
This paper presents new tokenization resources for Irish and evaluation measures of alignment with the morphological boundaries of the language. We present MoirfEolas, a dataset of over 35,000 Irish words mapped to their respective eclipses, prefixes and suffixes as well as an evaluation metric CríochScore, that evaluates the alignment of tokenizations with the morphological boundaries present in MoirfEolas. We evaluate common tokenization algorithms using CríochScore as well as intrinsic metrics present in the tokenization literature. We find that the Unigram Language Model aligns with Irish morphology more often than the other algorithms evaluated. We also find trade-offs between morphological-alignment of tokenization with both compression as well as vocabulary efficiency, providing practical insights for Irish natural language processing development. This dataset contributes towards combating the Irish language's low-resource status; moreover, the construction process reported in this paper can be emulated by other languages to create specialised morphological resources.
发表机构
- ADAPT Centre, Dublin City University(都柏林城市大学ADAPT中心)
- Trinity College Dublin(都柏林圣三一学院)
机构由 AI 辅助整理,请以论文原文为准。