发表机构
University of Kansas; University of Kansas Medical Center(堪萨斯大学; 堪萨斯大学医学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SoftGene提出一种基于蛋白质语言模型ESM的层次注意力编码器,结合软提示与硬提示的混合方案,用于可解释的基因集注释,并在GO和MSigDB基准上验证了其有效性。
AI 中文摘要
基因集分析是功能基因组学的基石,但它仍然劳动密集且严重依赖人工管理和专家生物学解释。尽管大型语言模型(LLMs)已成为基因组推理和注释的强大工具,但大多数现有方法依赖于符号基因名称,未能捕获领域特定的生物学结构,特别是控制分子活性、相互作用和下游基因功能的蛋白质序列信息。在这项工作中,我们提出了SoftGene,一种基于LLM的基因集注释的新框架,利用基因集的层次结构。首先,我们使用基于ESM(一种蛋白质语言模型)的层次注意力编码器,利用蛋白质水平的氨基酸序列信息来表示每个基因集。其次,我们构建了一种混合提示方案,将源自基因集嵌入的软提示与包含由LLM生成的辅助上下文的硬提示相结合,并将所得提示输入本地LLM进行注释。我们在两个基准数据集上评估了我们的框架:基因本体论(GO)和分子签名数据库(MSigDB)。我们的结果表明,将蛋白质序列表示与文本上下文相结合可整体改善基因集注释,而按领域分析则揭示蛋白质嵌入的贡献在不同生物学领域间存在差异。
英文摘要
Gene set analysis is a cornerstone of functional genomics, yet it remains labor-intensive and heavily dependent on manual curation and expert biological interpretation. While Large Language Models (LLMs) have emerged as powerful tools for genomic reasoning and annotation, most existing approaches rely on symbolic gene names and fail to capture domain-specific biological structure, particularly protein sequence information that governs molecular activity, interactions, and downstream gene function. In this work, we propose SoftGene, a novel framework for LLM-based gene set annotation that leverages the hierarchical structure of gene sets. First, we use a hierarchical attention-based encoder built on ESM, a protein language model, to represent each gene set using protein-level amino acid sequence information. Second, we construct a hybrid prompting scheme that combines soft prompts derived from gene set embeddings with hard prompts containing auxiliary context generated by an LLM, and feed the resulting prompt into a local LLM for annotation. We evaluate our framework on two benchmark datasets: Gene Ontology (GO) and the Molecular Signatures Database (MSigDB). Our results show that integrating protein-sequence representations with textual context improves gene set annotation overall, while per-domain analyses reveal that the contribution of protein embeddings varies across biological domains.
CommentsAccepted to EMNLP 2026 Main Conference