AI 中文总结
研究蛋白质结构建模中生成与理解的对齐问题,通过基准测试发现现有生成模型表现不佳,受REPA框架启发提出训练时使生成模型与理解模型对齐,实验表明该方法显著提升功能蛋白质生成。
AI 中文摘要
在深度神经网络训练中,理解和生成常被视为两个独立范式,尽管它们训练目标相关。此前研究表明视觉领域生成模型在理解任务中表现不佳,蛋白质领域情况未知。本文通过在蛋白质理解任务上对先进生成模型进行基准测试,发现其性能逊于现有蛋白质编码器。受表征对齐框架启发,提出训练中使生成性蛋白质扩散模型与预训练理解模型显式对齐。MotifBench实验表明,表征对齐显著提升功能蛋白质生成,如Protpardelle - 1c的MotifBench分数从39.2提高到47.1,相对提升20%。结果表明表征对齐为弥合蛋白质结构建模中理解与生成提供了通用有效机制。
英文摘要
Understanding and generation are often treated as two separate paradigms in training deep neural networks, despite the fact that both are trained with closely related objectives such as denoising and masked prediction. While prior studies have shown that generative models often learn suboptimal representations for understanding tasks in vision, it is less understood whether a similar gap exists in the protein domain. In this work, we systematically investigate this question by benchmarking state-of-the-art protein generative models on widely-used protein understanding tasks, and observe that these models exhibit consistently poor performance compared to existing protein encoders. Furthermore, inspired by the Representation Alignment (REPA) framework, we propose to explicitly align generative protein diffusion models with pretrained protein understanding models during training. Experiments on the MotifBench demonstrate that representation alignment significantly improves functional protein generation, boosting the MotifBench score of Protpardelle-1c from 39.2 to 47.1, corresponding to a 20% relative improvement. Our results suggest that representation alignment provides a general and effective mechanism for bridging understanding and generation in protein structure modeling.