发表机构
Yusuf Hamied Department of Chemistry, University of Cambridge; Department of Biotechnology, Agricultural University of Athens; Archimedes Unit, Athena Research Center(剑桥大学优素福·哈米德化学系; 雅典农业大学生物技术系; 阿瑞斯研究中心阿尔基米德单元)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MD-LLM-2是具备物理条件约束与显式路径概率的可迁移分子动力学语言模型,可低误差重现蛋白质构象特性,能预测未见过的力场参数的平衡种群及首次通过动力学。
AI 中文摘要
分子动力学(MD)模拟用于建模蛋白质在构象状态间的运动,但跨越自由能垒及对多种热力学条件进行采样的计算成本极高。我们近期推出了MD-LLM-1,该模型表明语言模型可更低成本地生成构象轨迹,但它针对每种蛋白质需单独建模,且未显式建模热力学或动力学特性。本文中我们推出MD-LLM-2,这是一种可迁移语言模型,可基于序列、温度和分子历史生成轨迹。由于每一步都定义了结构标记上的归一化分布,该模型为每条生成或模拟的轨迹分配了显式似然值。在mdCATH结构域上训练后,MD-LLM-2在12个未见过的结构域中有10个将回转半径中位数重现至MD的1埃范围内;预测Trp-cage在更高温度下的压缩程度降低;生成的L99A T4溶菌酶轨迹可达到已发表的过渡区主链结构。在丙氨酸二肽模拟中,基于扭转缩放参数λ进行条件约束,对未见过的耦合项的盆地种群预测平均绝对误差为0.09,而MBAR的误差为0.35;首次通过概率的误差为0.08,而可逆马尔可夫状态模型的误差为0.16。帧等熵采样进一步限制了自回归轨迹生成过程中的误差累积。综上,这些结果将MD-LLM从蛋白质特异性轨迹生成扩展至具备显式路径似然的可迁移、受物理条件约束的建模,表明该框架可预测训练期间未见过的力场参数的平衡种群,且初步显示其还可捕捉首次通过动力学特性。
英文摘要
Molecular dynamics (MD) simulations model the motion of proteins between conformational states, but crossing free energy barriers and sampling multiple thermodynamic conditions remain computationally costly. We recently introduced MD-LLM-1, which showed that language models can generate conformational trajectories at lower cost, but required a separate model for each protein and did not explicitly model thermodynamics or kinetics. Here we introduce MD-LLM-2, a transferable language model that generates trajectories conditioned on sequence, temperature and molecular history. Because each step defines a normalized distribution over structural tokens, the model assigns an explicit likelihood to every generated or simulated trajectory. Trained across mdCATH domains, MD-LLM-2 reproduces the median radius of gyration of MD within 1 Angstrom in 10 of 12 unseen domains, predicts reduced Trp-cage compaction at higher temperature, and generates L99A T4 lysozyme trajectories that reach published transition-region backbone structures. In alanine dipeptide simulations, conditioning on a torsional scaling parameter $λ$ predicts basin populations at unseen couplings with a mean absolute error of 0.09 versus 0.35 for MBAR, and first-passage probabilities with an error of 0.08 versus 0.16 for a reversible Markov state model. Frame-isoentropic sampling further limits the accumulation of errors during autoregressive trajectory generation. Together, these results extend MD-LLM from protein-specific trajectory generation to transferable, physically conditioned modelling with explicit path likelihoods, and show that the framework can predict equilibrium populations for force-field parameters not seen during training, with an initial indication that first-passage kinetics can also be captured.