En-ViMedNER:带有UMLS语义类型标注的英越双语生物医学语料库
En-ViMedNER: An English-Vietnamese Parallel Biomedical Corpus with UMLS Semantic Type Annotations
- College of Engineering and Computer Science, VinUniversity(VinUniversity工程与计算机科学学院)
- Faculty of Engineering and IT, University of Technology Sydney(悉尼科技大学工程与信息技术学院)
- Center for AI Research, VinUniversity(VinUniversity人工智能研究中心)
- Monash University, Australia(莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出首个带UMLS语义类型标注的英越双语生物医学NER语料库En-ViMedNER,经多环节构建后在两类NER任务中取得F1值结果,相关资源公开发布以推动越南语生物医学NLP研究。
AI中文摘要:
生物医学命名实体识别(NER)是医疗AI应用的基础,包括临床决策支持和医学信息抽取。尽管带有统一医学语言系统(UMLS)标注的语料库(如MedMentions)推动了英文生物医学NER的发展,但越南语尚无同类资源。本文提出En-ViMedNER,首个带有UMLS语义类型标注的英越双语生物医学NER语料库,UMLS语义类型是语言中立的代码,提供共享跨语言标签空间,确保与现有基于UMLS的资源直接可比。该语料库包含4392篇PubMed摘要对、44892组英越句子对,以及来自MedMentions ST21pv数据集适配的21种语义类型的202949组对齐实体提及对。为平衡质量与可扩展性,我们通过自动翻译、专家后编辑、大语言模型(LLM)辅助标签投影及人工验证与裁决构建该语料库。我们将En-ViMedNER描述为大规模银标准语料库,包含经人工审核和共识修正的迷你测试子集。我们在两种设置下评估En-ViMedNER:(i)越南语输入/越南语输出生物医学NER;(ii)英语输入/越南语输出跨语言NER。对于越南语NER,我们基准测试越南语监督编码器模型、英语监督多语言编码器模型及基于提示的LLM,最佳模型在测试集上F1值达52.70,在迷你测试集上达53.78。对于跨语言NER,我们基准测试编码器-解码器模型及基于提示的LLM,最佳模型在迷你测试集上F1值达45.44。我们公开发布该语料库、语料库构建流程及基线模型,以促进未来越南语生物医学自然语言处理研究。
英文摘要:
Biomedical Named Entity Recognition (NER) is fundamental to healthcare AI applications, including clinical decision support and medical information extraction. While corpora with Unified Medical Language System (UMLS) annotations, such as MedMentions, have driven progress in English biomedical NER, no comparable resource exists for Vietnamese. This paper presents En-ViMedNER, the first English-Vietnamese parallel biomedical NER corpus annotated with UMLS semantic types, which are language-neutral codes providing a shared cross-lingual label space and ensuring direct comparability with existing UMLS-based resources. The corpus contains 4,392 PubMed abstract pairs, 44,892 English-Vietnamese sentence pairs, and 202,949 aligned entity-mention pairs across 21 semantic types adapted from the MedMentions ST21pv dataset. To balance quality and scalability, we have constructed the corpus through automatic translation, expert post-editing, LLM-assisted label projection, and human verification and adjudication. We characterize En-ViMedNER as a large-scale silver-standard corpus with a human-audited and consensus-corrected mini-test subset. We evaluate En-ViMedNER in two settings: (i) Vietnamese-input/Vietnamese-output biomedical NER and (ii) English-input/Vietnamese-output cross-lingual NER. For Vietnamese NER, we benchmark Vietnamese-supervised encoder models, English-supervised multilingual encoder models, and prompt-based LLMs. The best model achieves an F1 score of 52.70 on the test set and 53.78 on the mini-test set. For cross-lingual NER, we benchmark encoder-decoder models and prompt-based LLMs. The best model achieves an F1 score of 45.44 on the mini-test set. We publicly release our corpus, corpus construction pipeline, and baseline models to facilitate future Vietnamese biomedical NLP research.