BioBigBird:一种用于处理生物医学文本中长程依赖的稀疏注意力模型
BioBigBird: A Sparse Attention Model for Long-Range Dependency Processing in Biomedical Text
浏览论文内容
中文总结 AI 辅助
本研究针对生物医学文本长程依赖处理需求,提出BioBigBird模型,采用稀疏注意力机制与多任务学习框架,在BLURB基准上取得与SOTA相当的结果,为专业领域长上下文语言模型开发提供有效方法。
中文摘要 AI 辅助
尽管面向特定领域的大型语言模型(LLM)已编码了海量生物医学知识,但其有限的上下文窗口往往会阻碍对文本内部及文本间细微关系的深度理解。为解决这一局限,我们提出BioBigBird,这是一种在大量生物医学文献和临床数据上预训练的双向语言模型,专门用于处理长程依赖。BioBigBird利用稀疏注意力机制处理最多4096个token的序列,其训练采用多阶段流程以缓解大规模预训练语料中的噪声。我们还通过多任务学习(MTL)框架提升其性能,该框架联合优化命名实体识别与关系抽取。在BLURB基准上的综合评估表明,经MTL增强的BioBigBird取得了与现有最优模型极具竞争力的结果。本研究为开发面向专业领域的强大长上下文语言模型提供了有效方法,证明了扩展序列处理对复杂文本分析的价值。我们的模型已公开于此httpsURL。
英文摘要
While domain-specific Large Language Models (LLMs) have encoded vast biomedical knowledge, their limited context windows often hinder a deep understanding of nuanced relationships within and across texts. To address this limitation, we introduce BioBigBird, a bidirectional language model pre-trained on extensive biomedical literature and clinical data, specifically designed to handle long-range dependencies. BioBigBird leverages a sparse attention mechanism to process sequences up to 4096 tokens, and its training incorporates a multi-stage process to mitigate noise from the large-scale pre-training corpus. We further enhance its performance by employing a multi-task learning (MTL) framework that jointly optimizes for Named Entity Recognition and Relation Extraction. Comprehensive evaluations on the BLURB benchmark reveal that our MTL-enhanced BioBigBird achieves highly competitive results against state-of-the-art models. Our work contributes an effective methodology for developing powerful, long-context language models for specialized domains, demonstrating the value of extended sequence processing for complex text analysis. Our models are publicly available at https://huggingface.co/collections/bisectgroup/biobigbird.
发表机构
- Wadhwani School of Data Science and AI(瓦德瓦尼数据科学与人工智能学院)
- Indian Institute of Technology Madras(印度马德拉斯理工学院)
机构由 AI 辅助整理,请以论文原文为准。