发表机构
Green University of Bangladesh(孟加拉国绿色大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究对孟加拉语医学NER进行了大规模基准测试,发现微调的XLM-RoBERTa以F1 0.5959创下新最先进水平,且多语言模型优于语言特定模型,提示仅流程效果远逊于微调模型。
AI 中文摘要
针对低资源语言的医学命名实体识别(NER)由于高度的语言变异性和领域特定标注语料库的稀缺性,仍然是一项具有挑战性的任务。本研究提出了一个全面的经验基准,评估了三种微调的Transformer编码器——BanglaBERT、多语言BERT(mBERT)和XLM-RoBERTa——与GPT-4o mini在零样本和少样本提示配置下对孟加拉语医学NER的性能。与先前仅在50个样本的有限子集上评估大型语言模型的研究不同,我们在完整的3,179个样本测试集上进行了大规模评估,提供了统计上稳健且可复现的基线。我们微调的XLM-RoBERTa模型达到了0.5959的F1分数,创造了新的最先进水平,超过了先前报道的0.5848的最佳结果。关键的是,我们证明了语言特定的BanglaBERT模型持续逊于其多语言对应模型,F1分数为0.4937,表明在高度专业化的临床环境中,预训练的领域多样性可能超过语言特异性。此外,我们对该任务进行了详细的逐实体类型分析,揭示Medicine和Specialist类别被高可靠性地识别,F1分数超过0.83,而Symptom类别尽管是最频繁的训练类别,仍然是最具挑战性的,F1分数为0.4367。最后,微调的Transformer模型比最优提示配置高出3.76倍,证实了仅提示的流程在低资源语言环境中的结构化临床实体提取方面仍然不足。
英文摘要
Medical Named Entity Recognition (NER) for low-resource languages remains a challenging task due to high linguistic variability and a scarcity of domain-specific annotated corpora. This work presents a comprehensive empirical benchmark evaluating three fine-tuned transformer encoders-BanglaBERT, multilingual BERT (mBERT), and XLM-RoBERTa-against GPT-4o mini under zero-shot and few-shot prompting configurations for Bangla medical NER. In contrast to prior studies that evaluated large language models on limited subsets of only 50 samples, we conduct a large-scale evaluation across the full test set of 3,179 samples, providing statistically robust and reproducible baselines. Our fine-tuned XLM-RoBERTa model achieves an F1- score of 0.5959, establishing a new state-of-the-art and surpassing the previously reported best result of 0.5848. Crucially, we demonstrate that the language-specific BanglaBERT model consistently underperforms its multilingual counterparts with an F1-score of 0.4937, indicating that pretraining domain diversity can outweigh language specificity in highly specialized clinical settings. Furthermore, we present a detailed per-entity-type analysis for this task, revealing that Medicine and Specialist categories are recognized with high reliability, achieving F1- scores above 0.83, while the Symptom category remains the most challenging with an F1-score of 0.4367 despite being the most frequent training class. Finally, fine-tuned transformer models outperform the optimal prompting configuration by a factor of 3.76, confirming that prompt-only pipelines remain inadequate for structured clinical entity extraction in low-resource language environments.
Comments6 pages, 2 figures. Accepted at the 2026 IEEE International Conference on Biomedical Engineering, Computer and Information Technology for Health (BECITHCON), Dhaka, Bangladesh