AI 中文总结
研究针对蛋白质突变响应全局表征缺失问题,提出 RegimeFormer 大型蛋白质扰动模型并配套 RegimeAtlas,其可识别蛋白质扰动 regime,改善取代预测,衍生先验优化下游建模,为蛋白质扰动景观研究提供可扩展框架。
AI 中文摘要
蛋白质语言模型可规模化地组织序列与结构,但目前仍缺乏蛋白质如何响应突变的全局表征。我们提出 RegimeFormer,这是一种大型蛋白质扰动模型,与 RegimeAtlas 配套构建,通过协调和索引生命之树中 202,556,313 条非冗余蛋白质序列实现。一个保留多样性的 100 万蛋白质子集提供高分辨率训练与推理层,其中 995,995 个蛋白质可生成 407,048,356 个残基的残基水平摘要,且按需提供取代特异性预测。在实验性深度突变扫描、分子基准测试、结构置信度及进化约束方面,RegimeFormer 可识别可重复的蛋白质水平扰动 regime,其组织残基脆弱性、适应性与预测不确定性。Regime 条件可改善取代特异性预测,在未见过的蛋白质、未见过的家族及低同源性评估下实现最大相对增益。RegimeFormer 衍生的分子先验进一步改善下游转录组与药物反应建模。综上,RegimeFormer 与 RegimeAtlas 提供了一个可扩展框架,用于绘制、预测和查询全局序列空间中的蛋白质扰动景观。
英文摘要
Protein language models organize sequence and structure at scale, but a global representation of how proteins respond to mutation remains lacking. We present RegimeFormer, a large protein perturbation model coupled to RegimeAtlas, constructed by harmonizing and indexing 202,556,313 non-redundant protein sequences across the tree of life. A diversity-preserving one-million-protein subset provides the high-resolution training and inference layer, with 995,995 proteins yielding residue-level summaries across 407,048,356 residues and substitution-specific predictions available on demand. Across experimental deep mutational scanning, molecular benchmarks, structural confidence and evolutionary constraint, RegimeFormer identifies reproducible protein-level perturbation regimes that organize residue fragility, adaptability and predictive uncertainty. Regime conditioning improves substitution-specific prediction, with the largest relative gains under unseen-protein, unseen-family and low-homology evaluation. RegimeFormer-derived molecular priors further improve downstream transcriptomic and drug-response modelling. Together, RegimeFormer and RegimeAtlas provide a scalable framework for mapping, predicting and querying protein perturbation landscapes across global sequence space.