发表机构
Beijing Key Laboratory of Brain-inspired Spiking Large Models, School of Computer Science, Peking University; Guangdong Provincial Key Laboratory of AI for Science and Computing (AI4SC), School of Artificial Intelligence for Science, Shenzhen Graduate School, Peking University; School of Electronic and Computer Engineering, Shenzhen Graduate School, Peking University(北京大学计算机学院; 北京大学深圳研究生院人工智能科学学院; 北京大学深圳研究生院电子与计算机工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RipplePLM,通过直接-远端交叉注意力解耦结构信息并引入性质潜在链注入生化监督,在MutaDescribe上显著提升突变效应生成的ROUGE-L。
AI 中文摘要
蛋白质突变效应生成任务要求模型用自然语言描述点突变的功能后果。现有的蛋白质到文本系统通常将突变信息编码为无差别的表示,忽视了突变诱导证据在结构和生化因素间的组织方式。我们提出RipplePLM,一种以直接-远端交叉注意力(DDCA)为核心的突变感知生成框架。通过从预训练蛋白质语言模型构建残基级别的突变扰动场,DDCA利用预测的接触图将突变表示组织为两条通路:突变位点的直接接触邻域及其多跳远端上下文。为补充这种结构解耦,我们进一步引入性质潜在链(PLChain),通过潜在性质标记将生化性质变化(如热稳定性和最适pH)的专家引导监督注入大语言模型的隐状态通路。在MutaDescribe上,RipplePLM在时间划分和结构划分上优于突变特异性基线;在匹配主干比较下,结构划分的平均ROUGE-L从{22.23}提升至{35.65}。专家评估进一步显示,与突变特异性基线相比,RipplePLM产生的生物学准确或相关描述比例更高。额外的消融实验、表示诊断和低N适应度回归实验进一步支持了所学突变感知表示的有效性。代码:此https URL。
英文摘要
Protein mutation effect generation asks a model to describe the functional consequence of a point mutation in natural language. Existing protein-to-text systems typically encode mutation information into undifferentiated representations, overlooking the organization of mutation-induced evidence across structural and biochemical factors. We propose RipplePLM, a mutation-aware generation framework centered on Direct-Distal Cross-Attention (DDCA). By constructing a residue-level Mutation Perturbation Field from pre-trained protein language models, DDCA leverages predicted contact maps to organize mutation representations into two pathways: the mutation site's immediate contact neighborhood and its multi-hop distal context. To complement this structural decomposition, we further introduce the Property Latent Chain (PLChain), which injects expert-guided supervision of biochemical property changes (e.g., thermostability and optimal pH) into the LLM hidden-state pathway through latent property tokens. On MutaDescribe, RipplePLM improves over mutation-specific baselines on temporal and structural splits; under a matched-backbone comparison, average structural-split ROUGE-L increases from 22.23 to 35.65. Expert evaluation further shows a higher proportion of biologically accurate or relevant descriptions than the mutation-specific baseline. Additional ablations, representation diagnostics, and low-$N$ fitness regression experiments further support the effectiveness of the learned mutation-aware representations. Code: https://github.com/Lyu6PosHao/RipplePLM.
CommentsNeurips 2026