arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAGenome:将基于检索的基因组语言模型扩展至长上下文

RAGenome: Scaling Retrieval-Based Genomic Language Models to Long Contexts

Frederikke Isa Marin, Panagiotis Antoniadis, Dionysia Danai Brilli, Andreas Bjerregaard, Rachael DeVries, Yan Li, Ole Winther, Wouter Boomsma

arXiv 2610.11761首次发表:更新:

发表机构

University of Copenhagen; Novo Nordisk A/S; Technical University of Denmark(哥本哈根大学; 诺和诺德公司; 丹麦技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出首个将预训练扩展至现有基于MSA的基因组语言模型100倍长上下文的基于检索的基因组语言模型RAGenome,其基因发现性能从0.45提升至0.60,训练成本低且兼具进化建模与长程能力。

AI 中文摘要

基因组包含调控细胞生物学特性的蓝图,因此深化对基因组功能的认知,对拓展生物学理解及推动生物医学进展均至关重要。大型语言模型在自然语言与蛋白质序列上的成功,促使研究者将类似思路应用于基因组数据。然而,标准基因组语言模型(gLMs)往往需要极大的计算资源,且在部分下游任务中仍落后于传统方法。近期,基于多序列比对(MSA)的预训练被提出作为高效替代方案,但现有模型受限于短输入上下文,仅能用于短程任务,如变异效应预测。本研究提出RAGenome,首个将预训练扩展至更长上下文(比现有基于MSA的gLMs长100倍)的基于检索的gLM,使其既能捕捉跨物种进化关系,又能捕捉物种内的长程相互作用。RAGenome在100种脊椎动物的全基因组比对数据上训练,大幅提升了基于MSA的gLMs的长程能力,将基因发现性能从0.45提升至0.60,同时在基于纯进化的任务(如致病变异优先级排序)上保持竞争力。RAGenome以极低的训练成本提供了有竞争力的gLM性能,在单一灵活可扩展框架内统一了进化建模与长程能力。代码可在该https URL获取。

英文摘要

The genome holds the blueprint that governs the biological properties of the cell. Consequently, advancing our knowledge of genomic function is crucial both for a broader understanding of biology and for continued biomedical advances. The success of large language models on natural language and protein sequences has motivated similar efforts on genomic data. However, standard genomic language models (gLMs) often require extremely large computational resources and still fall behind traditional methods on some downstream tasks. Recently, MSA-based pretraining has been proposed as an efficient alternative, but existing models are limited to short input contexts, restricting their use to short-range tasks, such as variant effect prediction. In this work, we present RAGenome, the first retrieval-based gLM that scales pretraining to longer contexts (100$\times$ longer than existing MSA-based gLMs), allowing it to capture both across-species evolutionary relationships and within-species longer-range interactions. Trained on whole-genome alignments from 100 vertebrates, RAGenome substantially improves the long-range capabilities of MSA-based gLMs, raising gene finding performance from 0.45 to 0.60, while remaining competitive on purely evolutionary-based tasks like prioritizing pathogenic variants. RAGenome provides competitive gLM performance at a fraction of the training cost, unifying evolutionary modeling and long-range capabilities within a single, flexible, scalable framework. Code is available at https://github.com/PanosAntoniadis/RAGenome.

Comments15 pages, 7 figures and 4 tables, Accepted at ML4Molecules workshop at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑