Assay-Aware BindingDB:为结合亲和力预测策展实验上下文
Assay-Aware BindingDB: Curating Experimental Context for Binding Affinity Prediction
浏览论文内容
中文总结 AI 辅助
针对生物活性数据异质性,提出Assay-Aware BindingDB数据集及上下文条件亲和力模型,通过两阶段智能体框架提取测定元数据,提升预测性能,验证了元数据的价值。
中文摘要 AI 辅助
蛋白质-配体结合亲和力预测是计算药物发现的基础,然而现代AI驱动模型受限于其训练数据中普遍存在的异质性:生物活性值跨不同测定类型和实验条件聚合,而未考虑方案级别的差异,引入了系统性噪声。现有的协调方法要么丢弃测定级别的元数据,要么将其压缩为粗糙的分类区分,留下丰富的上下文信号未被利用。我们通过两项贡献来解决这一差距。首先,我们引入了Assay-Aware BindingDB,通过一个两阶段智能体框架,从原始文献中提取结构化元数据,该框架将证据提取与本体条件化的JSON合成分离,增强了74,425个BindingDB蛋白质-配体对,涵盖四种测定类型(ITC、SPR、RBA和FPA)。与领域专家参考相比,该框架在所有测定类型中实现了存在性F1分数≥0.947和语义内容准确率≥0.913。其次,我们开发了一个上下文条件亲和力模型,将测定上下文嵌入注入Boltz-2亲和力模块。在论文级别的留出分割上,该模型将组合机制均方误差从1.27降至1.19,并将皮尔逊相关系数从0.64提升至0.67,其中SPR和FPA的增益显著。ITC是所考虑的唯一天然无标记且无固定化的测定,未显示改进,这与其设计消除了元数据所捕获的方案伪影一致。这些结果支持了系统性策展的测定元数据为亲和力预测提供信息性信号的假设。
英文摘要
Protein--ligand binding affinity prediction is fundamental to computational drug discovery, yet modern AI-driven models are limited by pervasive heterogeneity in their training data: bioactivity values are aggregated across diverse assay types and experimental conditions without accounting for protocol-level differences, introducing systematic noise. Existing harmonization approaches either discard assay-level metadata or collapse it into coarse categorical distinctions, leaving rich contextual signal unused. We address this gap with two contributions. First, we introduce Assay-Aware BindingDB, augmenting 74,425 BindingDB protein--ligand pairs across four assay types (ITC, SPR, RBA, and FPA) with structured metadata extracted from primary literature using a two-stage agentic framework that separates evidence extraction from ontology-conditioned JSON synthesis. Against domain-expert references, the framework achieves presence F1 $\geq 0.947$ and semantic content accuracy $\geq 0.913$ across all assay types. Second, we develop a context-conditioned affinity model that injects an assay-context embedding into the Boltz-2 affinity module. On a paper-level held-out split, the model reduces combined-regime MSE from $1.27$ to $1.19$ and increases Pearson correlation from $0.64$ to $0.67$, with significant gains for SPR and FPA. ITC, the only label- and immobilization-free assay considered, shows no improvement, consistent with its design removing the protocol artifacts the metadata captures. These results support the hypothesis that systematically curated assay metadata provides informative signal for affinity prediction.
发表机构
- The University of Texas Health Science Center at Houston(休斯顿德克萨斯大学健康科学中心)
- Texas A&M University(德克萨斯农工大学)
- The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)
- Mayo Clinic(梅奥诊所)
机构由 AI 辅助整理,请以论文原文为准。