发表机构
Trillion Labs; KAIST; SK Biopharmaceuticals Co., Ltd.; Lunit Inc.; AIGEN Sciences Inc.(Trillion Labs; 韩国科学技术院; SK生物制药公司; Lunit公司; AIGEN科学公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
为满足生物学大语言模型训练需求,提出TheBioCollection语料库,整合异构生物资源并丰富记录、引入新任务,配对TheBioCollection-Eval评估,固定架构训练后模型在各领域表现提升,语言能力基本不变。
AI 中文摘要
向生物学大语言模型(BioLM)的发展产生了对训练语料库的需求,以便赋予模型对生物学的真正理解。然而,现有的生物资源分散在异构格式中,未被组织成用于语言模型训练的连贯语料库。我们提出了TheBioCollection,一个526亿token的预训练规模语料库,将这些不同资源转换为统一的、可用于训练的形式。它不仅整合现有数据,还用工具计算的生物学特性丰富每条记录,并引入新的指令任务。我们将该语料库与TheBioCollection-Eval配对进行评估。在固定基础Gravity-16B-A3B架构的情况下,在TheBioCollection上训练使模型在TheBioCollection-Eval上的总分提高了一倍多,且各领域均有提升,同时基本保持一般语言能力不变。
英文摘要
The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with TheBioCollection-Eval, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on TheBioCollection more than doubles its overall score on TheBioCollection-Eval with gains in every domain, while leaving general linguistic ability nearly intact.