arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TheBioCollection:用于生物学的统一预训练规模语言模型语料库

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung

arXiv 2607.08803首次发表:更新:

发表机构

Trillion Labs; KAIST; SK Biopharmaceuticals Co., Ltd.; Lunit Inc.; AIGEN Sciences Inc.(Trillion Labs; 韩国科学技术院; SK生物制药公司; Lunit公司; AIGEN科学公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

为满足生物学大语言模型训练需求,提出TheBioCollection语料库,整合异构生物资源并丰富记录、引入新任务,配对TheBioCollection-Eval评估,固定架构训练后模型在各领域表现提升,语言能力基本不变。

AI 中文摘要

向生物学大语言模型(BioLM)的发展产生了对训练语料库的需求,以便赋予模型对生物学的真正理解。然而,现有的生物资源分散在异构格式中,未被组织成用于语言模型训练的连贯语料库。我们提出了TheBioCollection,一个526亿token的预训练规模语料库,将这些不同资源转换为统一的、可用于训练的形式。它不仅整合现有数据,还用工具计算的生物学特性丰富每条记录,并引入新的指令任务。我们将该语料库与TheBioCollection-Eval配对进行评估。在固定基础Gravity-16B-A3B架构的情况下,在TheBioCollection上训练使模型在TheBioCollection-Eval上的总分提高了一倍多,且各领域均有提升,同时基本保持一般语言能力不变。

英文摘要

The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with TheBioCollection-Eval, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on TheBioCollection more than doubles its overall score on TheBioCollection-Eval with gains in every domain, while leaving general linguistic ability nearly intact.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑