发表机构
Mohamed bin Zayed University of Artificial Intelligence; Carnegie Mellon University(穆罕默德·本·扎耶德人工智能大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CausalBind,一种基于V-结构因果模型和稀疏交互先验的蛋白质-分子虚拟筛选方法,通过理论分析保证概念可识别性,并在DUD-E和LIT-PCBA基准上超越强基线,实现跨靶标和骨架的泛化。
AI 中文摘要
蛋白质-分子虚拟筛选日益被表述为在共享嵌入空间中的表示学习问题。现有方法依赖于密集的整体对齐,将不变的结合决定因素与干扰相关性纠缠在一起,并限制了向新靶标的迁移。已有研究指出,蛋白质-分子系统中的结合涉及稀疏的跨模态交互:结合由较小的接触界面和少数决定性的局部交互(如氢键、疏水接触和盐桥)所控制,而非蛋白质和分子的整体结构。我们假设,发现并利用稀疏交互模式对于超越训练数据的泛化至关重要,因为这些模式是可复用的,并有望在不同场景中提升性能。在本文中,我们旨在识别并利用稀疏交互模式,并验证我们的假设。由于训练数据仅包含观察到的结合对,我们通过Heckman式选择下的V-结构因果模型形式化这一先验,并建立了三个理论结果:(i)在没有适当稀疏约束的情况下,相互作用的蛋白质和分子的潜在概念不可识别;(ii)在结构稀疏条件下,这些概念及其稀疏交互可逐分量识别;(iii)这些条件的低秩松弛产生概念和交互的子空间可识别性。受这些原理启发,我们提出了CausalBind及其三种实现变体。在DUD-E和LIT-PCBA基准上的大量实验表明,所有变体均持续优于强检索基线,在LIT-PCBA早期富集上提升最大,并进一步泛化到靶标级和骨架级的分布外划分。代码可在以下网址获取:此https URL。
英文摘要
Protein-molecule virtual screening is increasingly cast as a problem of representation learning in a shared embedding space. Existing methods rely on dense holistic alignment, entangling invariant binding determinants with nuisance correlations and limiting transfer to new targets. It has been noted that binding in protein-molecule systems involves sparse cross-modality interactions: binding is governed by a small contact interface and a few decisive local interactions (e.g., hydrogen bonds, hydrophobic contacts, and salt bridges) rather than the global structures of the protein and molecule. We hypothesize that uncovering and leveraging sparse interaction patterns is critical for generalization beyond the training data, as these patterns are reusable and expected to improve performance across different scenarios. In this paper, we aim to identify and leverage sparse interaction patterns, and verify our hypothesis. Since the training data contain only observed binding pairs, we formalize this prior via a V-structure causal model under Heckman-style selection, and establish three theoretical results: (i) the latent concepts of interacting proteins and molecules are not identifiable without appropriate sparsity constraints; (ii) these concepts and their sparse interactions are component-wise identifiable under structural sparsity conditions; and (iii) a low-rank relaxation of these conditions yields subspace identifiability of the concepts and interactions. Inspired by these principles, we propose CausalBind with three implementation variants. Extensive experiments on DUD-E and LIT-PCBA benchmarks show that all variants consistently outperform strong retrieval baselines, with the largest gains on LIT-PCBA early enrichment, and further generalize to target- and scaffold-level out-of-distribution splits. Code is available at https://github.com/lokali/CausalBind.
CommentsNeurIPS 2026 (Oral)