发表机构
William & Mary; Baylor University; Southwest University; Hefei University of Technology; Florida Atlantic University(威廉与玛丽学院; 贝勒大学; 西南大学; 合肥工业大学; 佛罗里达大西洋大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ColdDDI是一个基于DrugBank构建的诊断基准,用于评估冷启动DDI预测模型是否真正利用药理学证据,发现微调LLM在介导物缺失时性能显著下降,而KG增强基线常忽略证据。
AI 中文摘要
冷启动药物-药物相互作用(DDI)预测旨在检验模型能否为没有训练时相互作用历史的药物识别具有临床意义的相互作用。现有基准大多报告聚合的边预测分数,留下了一个关键的评估问题:当模型接收到分子、文本或知识图谱(KG)证据时,它们是否真正利用了药理学上支持相互作用的证据?我们引入了ColdDDI,一个基于DrugBank 5.1.13构建的可重建诊断基准,包含1,900种已批准的小分子药物和565,731个阳性DDI对。ColdDDI评估具有零个、一个或两个未见药物的配对。它还根据相互作用是否改变药物暴露或药物效应,以及生物医学知识图谱是否包含可合理介导相互作用的共享酶、转运体或靶点,对每个相互作用进行注释。这些注释将证据可用性与预测依赖性区分开来。我们评估了八种传统DDI方法和13种LLM;对于开放权重LLM,我们测试了五种提示模式,并使用掩蔽、药物替换和通道敏感性指标来探测知识利用。ColdDDI揭示,在两种药物均未见的最难分割中,主要性能差异在于介导物的可用性。一个微调的1B LLM恢复了具有共享酶、转运体或靶点的相互作用的89-93%,但在没有此类介导物时仅恢复40-62%。更重要的是,KG提供的证据并不总是被使用;当共享介导物被掩蔽或破坏时,几种KG增强的基线变化很小,而微调的LLM对此干预反应强烈。因此,ColdDDI评估的是知识利用而非仅仅是知识获取,展示了冷启动DDI模型在何处依赖机制证据,以及在接收到证据后仍失败的地方。代码可在该HTTPS URL获取。
英文摘要
Cold-start drug-drug interaction (DDI) prediction tests whether models can identify clinically significant interactions for drugs without training-time interaction history. Existing benchmarks mostly report aggregate edge-prediction scores, leaving a key evaluation question unanswered: when models receive molecular, textual, or knowledge-graph (KG) evidence, do they actually use the evidence that pharmacologically supports the interaction? We introduce ColdDDI, a reconstructible diagnostic benchmark built from DrugBank 5.1.13, with 1,900 approved small-molecule drugs and 565,731 positive DDI pairs. ColdDDI evaluates pairs with zero, one, or two unseen drugs. It also annotates each interaction by whether it changes drug exposure or drug effect, and by whether the biomedical knowledge graph contains shared enzymes, transporters, or targets that can plausibly mediate the interaction. These annotations separate evidence availability from predictive dependence. We evaluate eight conventional DDI methods and 13 LLMs; for open-weight LLMs, we test five prompt patterns and use masking, drug replacement, and channel-sensitivity metrics to probe knowledge utilization. ColdDDI exposes that, in the hardest split where both drugs are unseen, the main performance divide is mediator availability. A fine-tuned 1B LLM recovers 89-93% of interactions with a shared enzyme, transporter, or target, but only 40-62% without such a mediator. More importantly, KG-provided evidence is not always used; several KG-augmented baselines change little when the shared mediator is masked or disrupted, whereas fine-tuned LLMs respond strongly to this intervention. Thus, ColdDDI evaluates knowledge utilization rather than knowledge access alone, showing where cold-start DDI models rely on mechanistic evidence and where they fail despite receiving it. Code is available at https://github.com/0217ljh/ColdDDI-NeurIPS2026.
CommentsAccepted at NeurIPS 2026 (poster). Code: https://github.com/0217ljh/ColdDDI-NeurIPS2026