AI 中文总结
本研究针对多参数蛋白质工程的挑战,提出采用多任务神经网络的贝叶斯参数化方法,经27个多参数蛋白质数据集的2592个模型对比,发现贝叶斯最后层模型性能最优,降维与独热编码可提升表现,为该领域确立了实用设计原则。
AI 中文摘要
同时对多种蛋白质特性进行工程改造仍是一项重大挑战。现有的基于机器学习的蛋白质工程流程通常会分别对各特性进行建模,无法捕捉它们之间的依赖关系与权衡。在此,我们系统评估了多任务神经网络上的贝叶斯参数化如何在稀缺、有噪声的实验数据下实现稳健的多参数蛋白质工程。我们整理了一套包含27个多参数蛋白质数据集的综合集合,随后在16种序列表示及降维方法(共2592个模型)上,对比了从低到全贝叶斯参数化的三种算法架构。贝叶斯最后层模型在整体准确率、泛化能力与校准度上表现最强,在70%的基准数据集上排名最优。降维方法使各架构的预测性能提升最高达42%,校准度提升最高达57%。值得注意的是,简单的独热编码在25%的基准数据集上取得了最优性能,尤其在数据集规模较大时表现突出。这些结果为可靠且数据高效的多参数蛋白质工程确立了实用设计原则。
英文摘要
Simultaneously engineering multiple protein properties remains a major challenge. Existing machine learning-based pipelines for protein engineering often model properties separately, failing to capture their dependencies and trade-offs. Here, we systematically evaluate how Bayesian parameterization on Multitask Neural Networks can enable robust simultaneous protein engineering under scarce, noisy experimental data. We curated a comprehensive set of 27 multiparameter protein datasets. Then, we compared three algorithm architectures spanning low to full Bayesian parameterization across 16 sequence representations and dimensionality reduction (2,592 models). Bayesian Last Layer models delivered the strongest overall accuracy, generalization, and calibration, ranking as the top-performing model on 70% of benchmark datasets. Dimensionality reduction improved predictive performance by up to 42% and enhanced calibration up to 57% across architectures. Notably, simple One-Hot encoding achieved top performance on 25% of benchmark datasets, particularly with larger datasets. These results establish practical design principles for reliable and data-efficient multiparameter protein engineering.