arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从文献中自动识别生物信息学软件命名实体

Automatic bioinformatic software named entity recognition from literature

Hao Xuan, Rithvij Pasupuleti, Ben Liu, Haishuo Sun, Jun Zhang, Zijun Yao, Cuncong Zhong

arXiv 2608.19201首次发表:更新:

AI 中文总结

本研究提出混合命名实体识别框架SNAIL,结合词汇与语义建模策略,在基准数据集及真实文献上识别生物信息学软件/数据库,性能优于现有方法,可用于大规模文献的资源元分析。

AI 中文摘要

生物信息学软件和数据库是现代生命科学研究的重要组成部分,但它们在科学文献中的提及往往不一致,难以大规模系统识别。缺乏全面且最新的生物信息学资源目录,阻碍了自动化生物医学知识提取和简化数据分析的工作。本文提出SNAIL,这是一种混合命名实体识别框架,旨在从生物医学文本中自动识别生物信息学软件和数据库(SW/DB)名称。SNAIL整合了互补的词汇和语义建模策略:词汇组件捕捉SW/DB名称特有的正字法模式和上下文线索;语义组件利用SciBERT等基于Transformer的语言模型生成的上下文嵌入,并结合显式的token掩码策略以增强面向实体的表示。通过混合流水线自动构建了大型训练语料库,该流水线整合了引文提示提取和大语言模型辅助蒸馏。在两个独立基准数据集和真实研究文章上的评估表明,SNAIL的性能显著优于现有方法,包括bioNerDS2等领域特定方法,以及ChatGPT、Gemini、Grok和Claude等通用大语言模型。将SNAIL应用于大规模文献分析进一步揭示了生物信息学子领域中不同的期刊层面偏好。这些结果表明,SNAIL为识别科学文本中的生物信息学资源提供了准确且可扩展的解决方案,并支持对工具使用和研究趋势进行系统元分析。

英文摘要

Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑