arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00099q-bio.GNcs.AI

LLMBDC:面向生物域的基因本体聚类语言模型

LLMBDC: Language Model for Biological Domains Oriented Clustering of Gene Ontology

Ximing Ran, Jie Xu, Peng Jin, Zhaohui Qin, Zhexing Wen, Jiaying Lu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对GO富集分析的术语冗余问题,提出无需训练的LLMBDC框架,利用LLM零样本推理聚类GO术语为生物域,在AD和FXS数据集上较基线方法显著提升了聚类性能。

中文摘要 AI 辅助

基因本体(Gene Ontology, GO)富集分析是将大规模基因组数据转化为生物学见解的基础工具,但通常会产生数百个冗余术语,掩盖了总体主题。现有的汇总工具依赖固定相似度指标(如REVIGO、GOSemSim、clusterProfiler::simplify())、基因重叠度量(如Metascape)或静态层次映射(如GO-slim),因此无法融入生物学背景。人工整理可提供背景感知的分组,但具有主观性且劳动强度大。需要一种可扩展的、背景感知的框架,将GO术语聚类为可解释的高阶生物域。本文提出LLMBDC(Large Language Model for Biological Domains Oriented Clustering of Gene Ontology),这是一种无需训练的框架,利用大型语言模型(LLM)的零样本语义推理及置信度评分,仅通过推理时的本体信息将GO术语聚类为生物域(BioDomains)。在阿尔茨海默病(Alzheimer's disease, AD)和脆性X综合征(Fragile X syndrome, FXS)上针对包括SapBERT在内的6种基线方法进行基准测试,LLMBDC实现了显著更高的精度、召回率和聚类性能。与真实标注相比,LLMBDC在AD上的调整兰德指数(Adjusted Rand Index, ARI)较REVIGO从9.7%提升至73.3%,标准化互信息(Normalized Mutual Information, NMI)从59.9%提升至73.4%;在FXS上ARI从15.7%提升至66.6%,NMI从66.0%提升至79.5%。柯西组合检验进一步证实,聚合后的生物域保留了具有统计学意义的功能信号。LLMBDC为GO富集结果的背景感知、系统层面解读提供了一种可扩展、可复现且可解释的途径,同时保留了生物学特异性。

英文摘要

Gene Ontology (GO) enrichment analysis is a foundational tool for translating large-scale genomic data into biological insights, but typically yields hundreds of redundant terms that obscure overarching themes. Existing summarization tools rely on fixed similarity metrics (REVIGO, GOSemSim, clusterProfiler::simplify()), gene-overlap measures (Metascape), or static hierarchy mappings (GO-slim), and therefore cannot incorporate biological context. Manual curation provides context-aware grouping but is subjective and labor-intensive. A scalable, context-aware framework is needed to cluster GO terms into interpretable higher-order biological domains. Here we present LLMBDC (Large Language Model for Biological Domains Oriented Clustering of Gene Ontology), a training-free framework that leverages zero-shot semantic reasoning of LLMs with confidence scoring to cluster GO terms into BioDomains using only ontology information at inference time. Benchmarked across Alzheimer's disease (AD) and Fragile X syndrome (FXS) against six baseline methods including SapBERT, LLMBDC achieved substantially higher precision, recall, and clustering performance. Against ground-truth annotations, LLMBDC improved ARI from 9.7% to 73.3% (AD) and from 15.7% to 66.6% (FXS) over REVIGO, with corresponding NMI gains from 59.9% to 73.4% (AD) and 66.0% to 79.5% (FXS). A Cauchy combination test further confirmed that aggregated BioDomains retained statistically significant functional signals. LLMBDC provides a scalable, reproducible, and interpretable route to context-aware, system-level interpretation of GO enrichment results while preserving biological specificity.

发表机构

  • Emory University(埃默里大学)

机构由 AI 辅助整理,请以论文原文为准。

↑