发表机构
University of Cape Town(开普敦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低资源语言Runyankore,我们构建了首个公开NER基准RunyaNER,并系统研究辅助语言选择策略,发现基于嵌入的度量比传统语言特征更能预测迁移性能,为多语言迁移提供实用指导。
AI 中文摘要
跨语言零样本迁移和多语言微调是低资源语言中命名实体识别(NER)等NLP任务的有前景的方法,但在缺乏目标语言基准的情况下,尚不清楚哪种辅助语言选择策略能带来最佳迁移效果。我们推出了RunyaNER,这是首个公开可用的东非语言Runyankore的NER基准,并利用它来研究迁移时语言选择的问题。RunyaNER通过半自动化流程创建并经过完全人工验证,包含超过30,000个句子中的237,000多个标注词。我们在RunyaNER上对预训练模型进行了基准测试,证实我们的数据集具有足够的质量和规模,能够训练出有效的Runyankore NER模型。随后,我们利用RunyaNER研究跨语言零样本和多语言微调设置中的辅助语言选择。我们的实验表明,虽然迁移性能对辅助语言选择高度敏感,但基于标注训练跨度计算的嵌入度量与下游迁移性能的相关性,强于基于元数据或类型学的传统语言特征。通过发布RunyaNER并提供对辅助语言选择策略的系统分析,这项工作既贡献了新的基准资源,也为低资源环境下的多语言迁移提供了实用见解。
英文摘要
Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-resource languages, but in the absence of target language benchmarks, it is unclear which auxiliary language selection strategy leads to the best transfer. We introduce RunyaNER, the first publicly available NER benchmark for the East African language Runyankore, and use it to investigate the choice of which languages to use for transfer. Created with a semi-automated pipeline and fully manually verified, RunyaNER contains over 237k annotated words across 30k sentences. We benchmark pretrained models on RunyaNER, establishing that our dataset is of sufficient quality and size to produce effective Runyankore NER models. We then use RunyaNER to investigate auxiliary language selection in cross-lingual zero-shot and multilingual fine-tuning settings. Our experiments show that while transfer performance is highly sensitive to auxiliary language selection, embedding-based measures computed from labelled training spans correlate more strongly with downstream transfer performance than traditional linguistic features based on metadata or typology. By releasing RunyaNER and providing a systematic analysis of auxiliary language selection strategies, this work contributes both a new benchmark resource and practical insights for multilingual transfer in low-resource settings.
CommentsAccepted to the 6th Workshop on Multilingual Representation Learning (MRL 2026) at EMNLP 2026. Camera-ready version. 4 figures