arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14070q-bio.GNcs.LG

使用Evo 2探针筛选宏基因组数据中的生物安全特征

Screening of Biosecurity Features in Metagenomic Data with Evo 2 Probes

Jeremy Guntoro, Alexander Dack, Dylan Danno, Michaela Jančovičová, Križan Jurinović, Vanessa Smilansky

首次发表
浏览论文内容

中文总结 AI 辅助

研究利用Evo 2探针在宏基因组数据中筛选生物安全特征,通过训练线性和注意力探针检测抗菌抗性及细菌毒力等,发现其在检测AMR等方面效果良好,可作为快速低成本首过检测层,明确了该方法的优势与局限。

中文摘要 AI 辅助

基因组基础模型如Evo 2能学习丰富的序列表示,但在生物安全筛选方面价值未充分探索。本文通过在冻结的Evo 2第26层激活上训练最小线性和注意力探针,研究这些表示中与生物安全相关信号的线性可及性。结果表明,探针能有效检测抗菌抗性(AMR),区分更细粒度的AMR药物类别子类别并与无关功能基因分离,细菌毒力也可解码但较弱。AMR探针在模拟短读上无需重新训练就有可比排名性能,还对比了与其他模型的情况。这些结果表明基于轻量级嵌入的探针可作为宏基因组生物监测的快速、低成本首过检测层,并明确了该方法的优缺点。

英文摘要

Genomic foundation models such as Evo 2 learn rich sequence representations, but their value for biosecurity screening is largely unexplored. We ask how much biosecurity-relevant signal is linearly accessible in these representations by training minimal linear and attention probes on frozen Evo 2 layer-26 activations, without fine-tuning the underlying model. Across held-out metagenomic test sets, the probes detect antimicrobial resistance (AMR) with strong discrimination: a linear probe reaches a region-level ROC-AUC of 0.888 (mean-pool), rising to 0.977 with a single-head attention probe. The probes resolve finer-grained AMR drug-class subcategories and separate them from unrelated functional genes, providing additional evidence that the learned signal is not explained solely by generic functional-gene status. Bacterial virulence is also decodable, though more weakly (region-level ROC-AUC 0.833). The AMR probe retains comparable ranking performance on simulated short reads without retraining, enabling evaluation before assembly in settings where assembly is computationally costly or unreliable. It achieves a read-level ROC-AUC of 0.898 (mean-pool), comparable to the mean-pooled full-region result. Within SynGenome, AMR-associated prompt labels are only weakly recoverable from Evo 1.5-generated sequences; these prompt-derived labels do not establish the function of the generated response sequences. A complementary sparse-autoencoder analysis recovers interpretable resistance-associated features but proves less consistent than the supervised probes. Together, these results position lightweight embedding-based probes as a fast, inexpensive first-pass detection layer for metagenomic biosurveillance and map both strengths and current limits of the approach. This work was conducted as part of the AIxBio Hackathon 2026 hosted by BlueDot Impact, Apart Research, and Cambridge Biosecurity Hub.

发表机构

  • Department of Bioengineering, Imperial College London(帝国理工学院生物工程系)
  • John Innes Centre(约翰·英尼斯研究中心)
  • Independent Research Scientist(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

↑