发表机构
University of Arkansas at Little Rock(阿肯色大学小石城分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出DRET范式,将大型生物医学模型知识注入小型通用模型,在EBM-NLP语料库的PICO分类任务中,使66M参数的DistilBERT性能接近大10倍的模型且保持高效,可用于自动化文献综述等场景。
AI 中文摘要
BioBERT、ClinicalBERT等大型领域特定语言模型在生物医学NLP任务上表现出色,但它们的计算需求使其难以在许多实际部署场景中应用。DistilBERT等通用参数高效模型虽轻量,却缺乏PICO(人群、干预措施、对照、结局)分类等专业任务所需的领域知识。本文提出蒸馏快速嵌入迁移(Distilled Rapid Embedding Transfer,DRET),这是一种知识迁移范式,可将大型专业模型的生物医学领域知识注入小型通用模型,无需在原始专业语料库上重新训练。DRET是迭代策略系列:统一分词器合并策略(DRET 1.x)、混合嵌入平均法(DRET 2.0)、分层选择最权威源模型嵌入的基于优先级的嵌入迁移机制(DRET 3.x),结合嵌入层冻结、差分学习率、标签传播及感知类别不平衡的损失函数(DRET 4.x)。我们在类别严重不平衡的EBM-NLP语料库上,通过12项指标评估DRET在标记级PICO分类任务中的表现。经DRET增强的DistilBERT(6600万参数)的平衡准确率、召回率和ROC-AUC可与大一个数量级的模型相媲美,且在多项类别级指标上表现更优,同时保留了DistilBERT的效率。我们还通过余弦相似度、语义偏移和t-SNE分析证实迁移发生在嵌入层面。DRET为生物医学文本挖掘提供了可扩展、资源高效的途径,可实现接近领域专家的性能,直接应用于自动化系统文献综述和临床决策支持。
英文摘要
Large domain-specific language models such as BioBERT and ClinicalBERT achieve strong performance on biomedical NLP tasks, but their computational demands make them impractical for many real-world deployments. General-purpose, parameter-efficient models such as DistilBERT are lightweight yet lack the domain knowledge required for specialized tasks such as PICO (Population, Intervention, Comparison, Outcome) classification. We introduce Distilled Rapid Embedding Transfer (DRET), a knowledge-transfer paradigm that injects biomedical domain knowledge from large specialized models into a smaller general-purpose model without retraining on the original specialized corpora. DRET is developed as an iterative family of strategies: a unified tokenizer-merge strategy (DRET 1.x), hybrid embedding averaging (DRET 2.0), and a priority-based embedding-transfer mechanism (DRET 3.x) that hierarchically selects embeddings from the most authoritative source models, further combined with embedding-layer freezing, differential learning rates, label propagation, and imbalance-aware loss functions (DRET 4.x). We evaluate DRET on token-level PICO classification using the EBM-NLP corpus under severe class imbalance, across a twelve-metric battery. DRET-enhanced DistilBERT (66M parameters) attains balanced accuracy, recall, and ROC-AUC competitive with, and on several class-wise metrics exceeding, models an order of magnitude larger, while retaining DistilBERT's efficiency. We further show that transfer occurs at the embedding level through cosine-similarity, semantic-shift, and t-SNE analyses. DRET offers a scalable, resource-efficient route to near-domain-expert performance for biomedical text mining, with direct application to automated systematic literature reviews and clinical decision support.
Comments11 pages, 5 figures, 6 tables