AI 中文总结
本文针对生物医学变异-表型关系抽取,提出微调小型BERT模型(如DeBERTa)和Gemini Pro 1.0,在SNPPhenA语料库上接近或超越现有最先进水平。
AI 中文摘要
下一代测序技术彻底改变了对基因突变的研究,使得大规模探究其在疾病发展中的作用成为可能。然而,从海量生物医学文献中提取有意义的见解仍然是一个复杂挑战,无法通过人工方式解决。在本文中,我们提出了用于从生物医学文本中自动抽取关系的预训练模型(PTMs),特别针对变异-表型领域。我们在SNPPhenA语料库上的评估表明,微调基于BERT的小型模型(尤其是DeBERTa)能取得强劲性能,接近当前最先进水平(SOTA)。此外,我们的结果表明,仔细微调谷歌的Gemini Pro 1.0在句子级任务(模型仅处理目标句子)和摘要级任务(模型处理整个摘要)上均优于现有SOTA。
英文摘要
Next-Generation Sequencing has revolutionized the study of genetic mutations, enabling large-scale investigations into their roles in disease development. However, extracting meaningful insights from the vast amount of biomedical literature remains a complex challenge that cannot be addressed manually. In this paper, we present pre-trained models (PTMs) for the automatic extraction of relations from biomedical text, specifically targeting the variant-phenotype domain. Our evaluation on the SNPPhenA corpus demonstrates that fine-tuning small BERT-based models, particularly DeBERTa, yields strong performance, approaching the current state-of-the-art (SOTA). Additionally, our results indicate that carefully fine-tuning Google's Gemini Pro 1.0 outperforms the existing SOTA for both sentence-level tasks (where the model processes only the target sentence) and abstract-level tasks (where the model processes the entire abstract).