arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SeqMaestro:通过可解释机器学习从核苷酸序列到生物学假设

SeqMaestro: From nucleotide sequences to biological hypotheses through interpretable machine learning

Evgeny S. Saveliev, Krzysztof Kacprzyk, Charlotte Capitanchik, Neelanjan Mukherjee, Kate Matlin, Ryan Sheridan, Srinivas Ramachandran, Jernej Ule, David L. Bentley, Mihaela van der Schaar

arXiv 2609.14882首次发表:更新:

发表机构

University of Cambridge; The Francis Crick Institute; University of Colorado Anschutz Medical Campus(剑桥大学; 弗朗西斯·克里克研究所; 科罗拉多大学安舒茨医学校区)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SeqMaestro是一个无代码机器学习框架,通过两层接口连接核苷酸序列与可解释模型,利用多种模型组合识别稳健生物学信号,从而从序列数据提出生物学假设。

AI 中文摘要

核苷酸序列分析是调控基因组学、进化生物学和表型预测等问题的核心。经典生物信息学方法提取可解释的序列特征,如基序和k-mer组成,但其灵活性有限。相比之下,现代深度学习模型可以直接从原始序列中学习强大的预测表示,然而其内部表示和决策机制难以检查。可解释机器学习方法(如稀疏线性模型和决策树)提供了预测关系的人类可理解表示,但并非设计用于直接处理核苷酸序列。在此,我们介绍SeqMaestro,一个使用可解释模型从核苷酸序列提出生物学假设的机器学习框架。我们的解决方案围绕一个两层接口展开,该接口将核苷酸序列与更广泛的可解释机器学习生态系统连接起来。SeqMaestro利用此接口拟合可解释模型、特征表示和提取策略的多种组合,利用透明模型间的变异性来识别稳健的生物学信号和比单独特征重要性更丰富的预测关系。该系统还支持数据转换与清洗、模型拟合、超参数调优、可靠性分析,并将结果综合为情境化的书面报告。通过无代码工作流提供这些能力,SeqMaestro旨在使可解释序列分析对无需广泛编程或机器学习专业知识的研究者也可用。因此,SeqMaestro提供了一条从核苷酸序列到生物学假设的可访问途径。

英文摘要

Nucleotide sequence analysis is central to problems spanning regulatory genomics, evolutionary biology, and phenotype prediction. Classical bioinformatics methods extract interpretable sequence properties such as motifs and k-mer composition, but their flexibility is limited. In contrast, modern deep learning models can learn powerful predictive representations directly from raw sequences, yet their internal representations and decision mechanisms are difficult to inspect. Interpretable machine learning methods (e.g., sparse linear models and decision trees) provide human-understandable representations of predictive relationships but are not designed to operate directly on nucleotide sequences. Here, we introduce SeqMaestro, a machine learning framework that proposes biological hypotheses from nucleotide sequences using interpretable models. Our solution is centered around a two-layer interface that connects nucleotide sequences with the broader ecosystem of interpretable machine learning. SeqMaestro uses this interface to fit diverse combinations of interpretable models, feature representations, and extraction strategies, leveraging variability across transparent models to identify robust biological signals and richer predictive relationships than feature importance alone can provide. The system also supports data transformation and cleaning, model fitting, hyperparameter tuning, reliability analysis, and synthesis of results into a contextualized written report. By providing these capabilities through a no-code workflow, SeqMaestro is designed to make interpretable sequence analysis accessible to researchers without requiring extensive programming or machine learning expertise. SeqMaestro thereby provides an accessible route from nucleotide sequences to biological hypotheses.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑