资源极度匮乏下的西夏文分词:结合传统词典与未标注文本
Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled Text
浏览论文内容
中文总结 AI 辅助
针对资源极度匮乏的西夏文分词问题,该研究结合传统词典与未标注文本,提出含可靠性校准词典格等的框架,借助TangutEncoder在分词级五折交叉验证中F1达0.911,实现了超出有限监督词汇的泛化。
中文摘要 AI 辅助
西夏文是一种已灭绝的语言,其文字未明确标记词边界。我们开展了首个西夏文分词系统研究,使用2750个专家标注的分词(31893个词元)、传统词典和未标注文本。我们的框架结合可靠性校准的词典格表示、显式分布统计以及用MLM预训练的轻量字符编码器。分词级五折交叉验证显示,词汇和统计特征将CRF的F1值提升至约0.91;完整的TangutEncoder达到最高平均F1值0.911,且在已标注训练词汇之外提升了召回率。这些结果表明,模型在主题多样的保留段落上实现了有限监督词汇之外的泛化,而文档级迁移仍有待评估。
英文摘要
Tangut is an extinct language whose script does not explicitly mark word boundaries. We present the first systematic study of Tangut word segmentation using 2,750 expert-annotated segments (31,893 tokens), traditional lexicons, and unlabeled text. Our framework combines a reliability-calibrated lexicon-lattice representation, explicit distributional statistics, and a lightweight character encoder pretrained with MLM. In within-source five-fold cross-validation, the model integrating TangutEncoder, CRF, and external features obtains the numerically highest main-system mean F1 of 0.911 and substantially improves recall beyond the labeled training vocabulary. We further evaluate document-level transfer on 479 segments (4081 tokens) from five works absent from the annotated training corpus. You can access our project at https://github.com/jiangli-va/TangutSeg.
发表机构
- Chinese Academy of Social Sciences(中国社会科学院)
- University of Chinese Academy of Social Sciences(中国社会科学院大学)
- Rixin College, Tsinghua University(清华大学日新书院)
- Institute of Linguistics, Chinese Academy of Social Sciences(中国社会科学院语言研究所)
- Institute of Ethnology and Anthropology, Chinese Academy of Social Sciences(中国社会科学院民族学与人类学研究所)
- School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)
机构由 AI 辅助整理,请以论文原文为准。