发表机构
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SPINET利用分子动力学轨迹和细胞层表示,实现基于蛋白质运动的逆折叠,在mdCATH和ATLAS上优于静态和集成基线,显著提升序列恢复率。
AI 中文摘要
蛋白质在功能行使过程中会改变形状,然而大多数逆折叠模型是从单一的、固定的主链预测氨基酸序列。蛋白质工程中的一个核心挑战是设计经历特定运动的蛋白质,这需要考虑其结构随时间的变化。这促使了基于蛋白质运动的逆折叠。我们引入了SPINET,它从分子动力学轨迹中预测序列。它使用细胞层(cellular sheaves)来表示每一帧内的残基相互作用,并使用循环单元来整合跨帧的信息,然后一次性预测所有氨基酸。我们在mdCATH和ATLAS上评估了SPINET,在序列恢复方面,它优于所有评估的静态和集成基线。在mdCATH上,它实现了56.7%的top-1恢复率,而最强的静态基线为44.5%,最强的集成基线为40.7%。我们还评估了预测序列是否与沿目标轨迹采样的构象兼容。在mdCATH上,它们的中位TM-score为0.760,并且对于99.5%的测试域,结构恢复偏向于目标构象而非无关的诱饵。
英文摘要
Proteins change shape as they function, yet most inverse folding models predict amino acid sequences from a single, fixed backbone. A central challenge in protein engineering is to design proteins that undergo specific motions, which requires accounting for how their structures change over time. This motivates inverse protein folding conditioned on protein motion. We introduce SPINET, which predicts sequences from molecular dynamics trajectories. It uses cellular sheaves to represent residue interactions within each frame and recurrent units to integrate information across frames, then predicts all amino acids in a single pass. We evaluate SPINET on mdCATH and ATLAS, where it outperforms all evaluated static and ensemble baselines in sequence recovery. On mdCATH, it achieves 56.7% top-1 recovery, compared with 44.5% for the strongest static baseline and 40.7% for the strongest ensemble baseline. We also evaluate whether the predicted sequences are compatible with conformations sampled along the target trajectory. On mdCATH, they achieve a median TM-score of 0.760, and structural recovery favors target conformations over unrelated decoys for 99.5% of test domains.