arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LEMON-ZEST:进化信息感知的分词策略用于高效蛋白质语言建模

LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling

Biswajit Banerjee, Claudia Alvarez Carreno, Anton S. Petrov

arXiv 2609.37675首次发表:更新:

发表机构

University College London; Georgia Institute of Technology(伦敦大学学院; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出进化信息感知的分词方法ZEST及紧凑模型LEMON,以200M参数超越600M-3B模型,证明进化先验可替代大规模参数扩展,实现高效蛋白质表示学习。

AI 中文摘要

蛋白质语言模型(PLMs)在遵循自然语言处理中建立的缩放定律后,在基于序列和结构的任务上取得了显著进展,然而分词(tokenization)的潜力仍未得到充分利用。与人类语言不同,蛋白质在广泛的序列变异中仍保持结构,这是标准分词策略根本未能捕捉的特性。我们引入了ZEST(序列特征的区域编码,Zoned Encoding of Sequence Traits),这是一种从多序列比对保守区域中衍生的进化信息感知词汇表。ZEST允许在分词阶段直接嵌入域级生物学先验,而非通过规模隐式学习这些先验。ZEST原生地将序列压缩至平均每个token包含4个残基,使我们的模型能够在标准的1024-token上下文窗口中处理4000个残基。在此基础上,我们提出了LEMON(从自然中分层提取分子有序性,Layered Extraction of Molecular Ordering from Nature),这是一个紧凑的200M参数基于序列的模型,用于检测蛋白质序列之间的远程同源性,仅在单个H100 GPU上训练一周。尽管规模不大,LEMON在性能上超越了从600M到3B参数不等的最先进模型。我们的结果表明,进化信息感知的分词可以替代大规模参数扩展,为高效、基于生物学的蛋白质表示学习开辟了新方向。所有代码、模型权重和结果均在MIT许可证下公开提供。

英文摘要

Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and structure-based tasks, yet the potential of tokenization remains underexploited. Unlike human language, proteins preserve structure despite extensive sequence variation a property standard tokenization strategies fundamentally fail to capture. We introduce ZEST (Zoned Encoding of Sequence Traits), an evolution-informed vocabulary derived from conserved regions of multiple sequence alignments. ZEST allows embedding domain-level biological priors directly at the tokenization stage rather than learning them implicitly through scale. ZEST natively compresses sequences to an average token length of 4 residues, enabling our model to process 4,000 residues within a standard 1024-token context window. Building on this, we present LEMON (Layered Extraction of Molecular Ordering from Nature), a compact 200M-parameter sequence-based model for detection of remote homology between protein sequences trained on a single H100 GPU for one week. Despite its modest size, LEMON outperforms state-of-the-art models ranging from 600M to 3B parameters. Our results demonstrate that evolution-informed tokenization can substitute for massive parameter scaling, opening a new direction for efficient, biologically-grounded protein representation learning. All code, model weights, and results are publicly available under the MIT license.

CommentsAccepted to NeurIPS 2026. 9 pages, 4 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑