全球健康出版物的自定义命名实体识别与主题分类
Custom Named Entity Recognition and Topic Classification for Global Health Publications
- University of Geneva(日内瓦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对全球健康文献在资源受限下的NLP模型选择与调整问题,通过实验比较词向量、实体识别和主题分类方法,揭示了领域专业化与轻量调整的实用价值及Transformer高成本高精度的权衡,为知识系统构建提供实证依据。
AI中文摘要:
在标注数据和计算资源有限的环境中,应如何为全球健康文献选择和调整自然语言处理模型?本论文通过语义标签发现、命名实体识别(NER)和多标签主题分类的实验来研究这些挑战。首先,将在逐渐增大的专门语料库上训练的skip-gram word2vec模型与BioWordVec进行比较,以评估语料库大小和领域背景对标签发现的影响。词汇覆盖率和定性评估表明,更广泛的覆盖率不一定能产生更有用的领域特定关联。分析随后转向实体提取,将卷积spaCy模型与基于RoBERTa的Transformer在1,000个标注句子上进行比较。在宽松的评分协议下,Transformer达到0.80的微平均F1分数,而卷积模型为0.65-0.69,但前者耗时82秒而非5-6秒。这种权衡促使对卷积模型进行微调,并集成一个在NCBI疾病语料库上达到81.33%测试F1分数的疾病识别器。结合PDF预处理、实体过滤和MeSH丰富化,所得流程支持文档级索引。为了用主题标注补充实体提取,将基于MiniLM的少样本分类与BART-MNLI零样本推理在50个主题和1,000个手工构建的测试句子上进行比较。BART-MNLI达到95.2%的单标签准确率,而后者为59%;在部分人工评估下,报告的多标签准确率分别为88%和32%。然而,其较高的推理成本限制了实际集成。结果表明,领域专业化和轻量级调整在何处提供实用价值,以及Transformer的准确性在何处证明较高的推理成本是合理的,为在资源约束下构建知识系统提供了实证基础。
英文摘要:
How should natural language processing models be selected and adapted for global health literature in environments where annotated data and computational resources are limited? This thesis investigates these challenges through experiments on semantic tag discovery, named entity recognition (NER), and multi-label topic classification. First, skip-gram word2vec models trained on progressively larger specialized corpora are compared with BioWordVec to assess how corpus size and domain context influence tag discovery. Vocabulary coverage and qualitative evaluation indicate that broader coverage does not necessarily yield more useful domain-specific associations. The analysis then turns to entity extraction, comparing convolutional spaCy models with a RoBERTa-based transformer on 1,000 annotated sentences. Under a lenient scoring protocol, the transformer achieves 0.80 micro-F1 versus 0.65-0.69 for convolutional models, but takes 82 seconds rather than 5-6 seconds. This trade-off motivates fine-tuning convolutional models and integrating a disease recognizer that achieves 81.33% test F1 on the NCBI Disease Corpus. Combined with PDF preprocessing, entity filtering, and MeSH enrichment, the resulting pipeline supports document-level indexing. To complement entity extraction with thematic annotation, MiniLM-based few-shot classification is compared with BART-MNLI zero-shot inference across 50 topics and 1,000 handcrafted test sentences. BART-MNLI achieves 95.2% single-label accuracy versus 59%; reported multi-label accuracies are 88% and 32% under partly manual assessment. However, its higher inference cost limits practical integration. The results show where domain specialization and lightweight adaptation offer practical value, and where transformer accuracy justifies higher inference costs, providing an empirical basis for building knowledge systems under resource constraints.