arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MiNER:针对临床文本中疟疾疾病实体识别的微调生物医学自然语言处理模型

MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts

V. S. Anoop, Devika N

arXiv 2609.00073首次发表:更新:

发表机构

Amrita Vishwa Vidyapeetham; Kerala University of Digital Sciences, Innovation and Technology(阿姆里塔维什瓦维迪亚皮塔姆大学; 喀拉拉邦数字科学、创新与技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出微调BioBERT的MiNER模型,从疟疾文献中提取生物医学信息,在精确率等指标上优于对比方法,并发布了人工标注的实体关系数据集供同行使用。

AI 中文摘要

疟疾仍是重大的全球卫生负担,需要持续开展研究以理解其复杂的分子机制、流行病学特征及潜在治疗干预手段。从海量且不断增长的疟疾文献中提取关键生物医学信息是一项具有挑战性的任务,需要创新方法。近年来,预训练语言模型彻底改变了自然语言处理任务,在多个领域展现出卓越能力。本文提出一种微调后的预训练生物医学语言模型,用于从疟疾疾病相关科学文献中提取生物医学信息。该方法首先选择并预处理大型疟疾科学文献语料库,然后用具有临床意义的实体对其进行标注;接着利用当前最先进的预训练语言模型BioBERT将文本数据编码为上下文感知表示;再通过领域特定标注和监督学习对模型进行微调,以增强其提取相关生物医学命名实体的能力。大量实验及与不同编码方法、机器学习算法的对比表明,所提方法在精确率、召回率和准确率上均显著优于对比方法;同时,本文还发布了人工标注的实体与关系提取数据集,以便其他卫生信息学研究者训练先进模型用于疟疾信息提取。

英文摘要

Malaria remains a significant global health burden, necessitating continuous research efforts to understand its complex molecular mechanisms, epidemiology, and potential therapeutic interventions. Extracting essential biomedical information from the vast and constantly growing malaria literature is a challenging task that demands innovative approaches. Recently, pre-trained language models have revolutionized natural language processing tasks, demonstrating remarkable capabilities in various domains. This paper proposes a fine-tuned pre-trained biomedical language model for biomedical information extraction from scientific literature on malaria disease. The proposed methodology selects and preprocesses a large corpus of scientific articles on malaria, and then annotates them with entities of clinical significance. It then leverages BioBERT, a state-of-the-art pre-trained language model, to encode the textual data into context-aware representations. We fine-tune the model using domain-specific annotations and supervised learning to enhance its ability to extract relevant biomedical named entities. Extensive experiments and comparisons with different encoding and machine learning algorithms show that the proposed approach significantly outperforms them in precision, recall, and accuracy. We also publish our human-labeled dataset for entity and relation extraction to enable other health informatics researchers to train advanced models for malaria information extraction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑