探测与操控Boltz-1主干-扩散边界处的生物学特性
Probing and steering biology across Boltz-1s trunk-diffusion boundary
浏览论文内容
中文总结 AI 辅助
该研究以Boltz-1为对象,分析其主干与扩散模块的生物学信息传递差异,验证线性可解码方向的操控性,发布相关资源。
中文摘要 AI 辅助
AlphaFold3类结构预测模型将处理序列与上下文的表征主干(trunk)与生成原子坐标的扩散模块(diffusion module)相结合。生物学信息跨越该架构边界时的变化机制仍鲜为人知。我们使用线性探针(linear probes)、稀疏自编码器(SAEs)和因果干预方法,分析Boltz-1的Pairformer主干与扩散模块的残基级激活情况。从主干中,几何特征(二级结构、无序性)和序列化学特征(氨基酸种类、信号肽、二硫键注释)均可线性解码;而在扩散模块中,二者出现分化:二级结构基本不变地传递,序列化学特征则大幅衰减。随后我们测试可解码方向是否能操控模型,对调控扩散模块的最终主干单表征进行干预:螺旋和线圈方向会使预测结构与匹配范数的随机对照呈剂量依赖性变化,但预测性极高(F1=0.82)的β-链方向未产生可测量的链含量增加,即线性可解码性并不意味着在我们测试的位点具有因果影响。与密集DSSP标签相比,相同探针对稀疏SwissProt注释的得分显著更低,因为模型正确预测的未注释残基被判定为假阳性,因此此类得分是下界。最后,在已有标签的位置,监督探针的表现优于单个SAE特征。我们发布了训练后的主干与扩散模块SAE、Boltz-1残基级激活数据及分析代码。
英文摘要
AlphaFold3-class structure predictors pair a representational trunk, which processes sequence and context, with a diffusion module, which generates atomic coordinates. How biological information changes as it crosses this architectural boundary remains poorly understood. We analyze per-residue activations from the Pairformer trunk and diffusion module of Boltz-1 using linear probes, sparse autoencoders (SAEs), and causal interventions. From the trunk, both geometry (secondary structure, disorder) and sequence chemistry (amino-acid identity, signal peptides, disulfide-bond annotations) are linearly decodable. In the diffusion module, the two diverge. Secondary structure transfers essentially unchanged, whereas sequence chemistry is strongly attenuated. We then test whether decodable directions can steer the model, intervening on the final trunk single representation that conditions the diffusion module. Helix and coil directions change predicted structure dose-dependently against matched-norm random controls, but a beta-strand direction that is highly predictive (F1 =0.82) produces no measurable increase in strand content: linear decodability does not imply causal influence at the site we tested. The same probes also score markedly lower against sparse SwissProt annotations than against dense DSSP labels, because unannotated residues that the model gets right are charged as false positives; such scores are therefore lower bounds. Finally, supervised probes outscore single SAE features wherever a label already exists. We release the trained trunk and diffusion SAEs, Boltz-1 per-residue activations, and the analysis code.
发表机构
- University of Oxford(牛津大学)
- Evolvere Biosciences(Evolvere生物科学公司)
- Botnar Research Centre(博特纳研究中心)
- Kavli Institute for Nanoscience Discovery, University of Oxford(牛津大学卡夫利纳米科学发现研究所)
- Department of Chemistry, University of Oxford(牛津大学化学系)
机构由 AI 辅助整理,请以论文原文为准。