MIRAGE:亲和力泛化中的插值与冗余度测量
MIRAGE: Measuring Interpolation and Redundancy in Affinity GEneralization
- DeepBio Scientific(DeepBio科学)
- University of Central Florida(中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对共折叠模型在亲和力预测中的评估缺陷,提出MIRAGE基准,通过家族支持轴区分可迁移原理与冗余记忆,揭示性能膨胀源于家族支持,并建议报告超额基线表现。
AI中文摘要:
深度学习现已支撑基于结构的药物设计,涵盖从复合物和亲和力预测到配体排序与构象生成。近期共折叠模型据称能以低得多的成本接近自由能微扰的精度。然而,标准评估——单一的留出集相关性或合并的构象成功率——无法将可迁移的结合原理与在公共数据库中反复接触相关蛋白质家族所导致的效应区分开来,而实际成功取决于真正的新靶标。我们提出MIRAGE(亲和力泛化中的插值与冗余度测量),一个即插即用的基准,将历史公共家族支持(截至2019年)作为显式变量,通过匹配分层、家族不相交对照、仅配体基线和时间评估,将家族支持轴应用于亲和力和构象预测。共折叠器的亲和力精度随家族支持显著上升,而无法利用测试家族的浅层对照保持平稳,共折叠器的提升幅度大,而每个家族不相交或平凡对照的提升接近于零。对于Nesso-1,该效应通过了协变量、条件、平衡和聚类检验;Boltz-2的终点受覆盖率限制。该效应定位于家族支持而非配体化学,仅从家族身份即可接近某一水平。在新家族上排名反转,家族不相交的随机森林在两个共折叠器上均领先,且相对于Nesso-1显著。在一个外部低支持靶标上,两个共折叠器均未超过分子量,这提供了佐证性而非总体性证据。gnina在重打分中显示出显著的家族支持依赖性,而smina则没有;无MSA的构象引擎比smina重对接显示出更大的差距。这种由冗余驱动的膨胀不同于传统的泄漏。我们建议报告跨家族支持的性能以及相对于支持不敏感基线的超额表现,并发布MIRAGE作为可安装的基准和数据集。
英文摘要:
Deep learning now underpins structure-based drug design, from complex and affinity prediction to ligand ranking and pose generation. Recent co-folding models reportedly approach free-energy-perturbation accuracy at far lower cost. Yet standard evaluation, a single held-out correlation or pooled pose-success rate, cannot separate transferable binding principles from repeated exposure to related protein families in public databases, and practical success depends on genuinely novel targets. We introduce MIRAGE (Measuring Interpolation and Redundancy in Affinity GEneralization), a plug-in benchmark treating historical public family support (through 2019) as an explicit variable, applying a family-support axis to affinity and pose prediction via matched strata, family-disjoint controls, ligand-only baselines, and temporal evaluation. Co-folder affinity accuracy rises sharply with family support, while shallow controls that cannot exploit the test family stay flat, large for co-folders and near zero for every family-disjoint or trivial control. For Nesso-1 it survives covariate, conditioning, balancing, and clustering checks; Boltz-2's endpoint is limited by coverage. It localizes to family support rather than ligand chemistry, approaching a level from family identity alone. Rankings reverse on novel families, where a family-disjoint random forest leads both co-folders, significantly vs Nesso-1. On one external low-support target, neither co-folder beats molecular weight, corroborative rather than population-level evidence. gnina shows significant support dependence in rescoring whereas smina does not; MSA-free pose engines show larger gaps than smina redocking. This redundancy-driven inflation differs from conventional leakage. We propose reporting performance across family support plus excess over a support-insensitive baseline, and release MIRAGE as an installable benchmark and dataset.