arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19071cs.CL

大型语言模型在生物医学关系抽取中的基准测试

Benchmarking Large Language Models for Biomedical Relation Extraction

  • Interdisciplinary School of Doctoral Studies(跨学科博士研究学院)
  • Faculty of Mathematics and Computer Science, University of Bucharest(布加勒斯特大学数学与计算机科学学院)
  • HLT Research Center(HLT研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Claudiu Creanga, Teodor Marchitan, Liviu P. Dinu

AI总结:

本研究在SNPPhenA语料库上基准测试多种LLM,发现OpenAI O1和微调Gemini 2.0 Pro分别在句子级和关联强度分类中取得最佳性能,证实了现代LLM在基因组知识提取中的优势。

AI中文摘要:

从生物医学文献中提取SNP-表型关联至关重要但具有挑战性。我们在SNPPhenA语料库上对多种NLP模型进行了基准测试,包括MLM、混合架构以及最先进的LLM(Gemini 2.0、OpenAI O系列、Qwen、Mistral),涵盖三个任务:句子级、摘要级和关联强度分类。OpenAI O1在非微调句子级分类中通过少样本学习取得了最先进(SOTA)结果(F1 0.89),并在摘要级分类中建立了新的SOTA(F1 0.82)。关联强度分类被证明是困难的,尽管在首次对该任务的LLM评估中,微调的Gemini 2.0 Pro表现最佳(F1 0.60)。专有LLM,尤其是在少样本(O1)或微调(Gemini 2.0 Pro)设置下,显著优于其他模型。这些发现证实了现代LLM在基因组知识提取中的强大能力。

英文摘要:

Extracting SNP-phenotype associations from biomedical literature is vital but challenging. We benchmarked diverse NLP models, including MLMs, hybrid architectures, and state-of-the-art LLMs (Gemini 2.0, OpenAI O-series, Qwen, Mistral), on the SNPPhenA corpus across three tasks: sentence-level, abstract-level, and association strength classification. OpenAI O1 achieved state-of-the-art (SOTA) results using few-shot learning for non-finetuned sentence-level classification (F1 0.89) and established a new SOTA for abstract-level classification (F1 0.82). Association strength classification proved difficult, though fine-tuned Gemini 2.0 Pro performed best (F1 0.60) in the first LLM evaluation of this task. Proprietary LLMs, especially in few-shot (O1) or fine-tuned (Gemini 2.0 Pro) settings, significantly outperformed other models. These findings confirm the power of modern LLMs for genomic knowledge extraction.

↑