AI 中文总结
该研究针对蛋白质结构生成模型的新颖性评估问题,引入DRR指标,分析8种模型的输出,提出零训练基线RetFold,成本远低于现有模型。
AI 中文摘要
蛋白质主链生成模型常被认为仅基于与已知蛋白质的低全链相似度就能探索新颖折叠空间,但这无法区分真正的新折叠与已知结构单元的新颖组装。我们首先探究这种粒度不匹配是否能单独解释已报道的比例,并引入结构域检索率(DRR),即生成的主链中任何组成结构域与CATH S40中已知结构域匹配的比例。将其应用于涵盖扩散和流匹配范式的8个主链生成模型,DRR在大多数输出中发现了可局部比对的已知结构,而包含被充分覆盖的完整结构域的比例则小得多,且取决于评分约定。为校准仅检索能实现的效果,我们提出RetFold,这是一种零训练基线方法,通过检索CATH结构域并基于几何的螺旋-连接优化来细化结构域间连接以构建主链,仅在CPU上的成本就低两个数量级。
英文摘要
Protein backbone generation models are often credited with exploring novel fold space based solely on low full-chain similarity to known proteins, yet this cannot distinguish a genuinely new fold from a novel assembly of known structural units. We first ask whether this granularity mismatch alone explains the reported rates, and introduce the Domain Retrieval Rate (DRR), the fraction of generated backbones for which any constituent domain matches a known domain in CATH S40. Applied to eight backbone generation models spanning diffusion and flow-matching paradigms, DRR finds locally alignable known structure in most outputs, while the fraction containing a substantially covered complete domain is considerably smaller and depends on the scoring convention. To calibrate what retrieval alone can achieve, we propose RetFold, a zero-training baseline that constructs backbones by retrieving CATH domains and refining inter-domain connections through geometry-based helix-linker optimization, at two orders of magnitude lower cost on CPU alone.