MolBioKG:通过多分辨率结构锚定将图外分子锚定到生物医学知识图谱中
MolBioKG: Grounding Out-of-Graph Molecules in Biomedical Knowledge Graphs via Multi-Resolution Structural Anchoring
AI总结:
MolBioKG是解决生物医学KG中图外分子冷启动问题的两层系统,通过多分辨率结构锚定实现分子与KG的关联,在多项任务中优于基线并提升关键指标。
AI中文摘要:
生物医学知识图谱(KG)可加速药物发现,但标准流程假设查询分子已作为图实体存在,导致未注册分子处于孤立状态。我们解决这一被称为图外分子问题的冷启动挑战,方法是引入MolBioKG。该两层系统通过多分辨率结构锚定将未见分子锚定到生物医学证据中,它将包含274万个分子(以骨架、片段、功能基团和指纹表示)的索引与包含960万条边的KG相连。仅给定SMILES字符串,MolBioKG即可检索结构相关的图实体并遍历其生物医学邻域,无需特定任务训练。它具备两种推理机制:使用Reciprocal Rank Fusion的静态多锚检索,以及用于自适应遍历的工具使用型LLM策略Adapt-KG。在图内链接恢复、复杂多跳推理和图外泛化任务的评估中,MolBioKG优于强基线。值得注意的是,它将多跳推理的Hits@10从0.585提升至0.876,将图外目标召回率从0.145提升至0.269,同时确保预测保留可追溯的结构锚点和来源可追溯的KG证据。
英文摘要:
Biomedical knowledge graphs (KGs) accelerate drug discovery, but standard pipelines assume query molecules already exist as graph entities, leaving unregistered molecules disconnected. We address this cold-start challenge, termed the out-of-graph molecule problem, by introducing MolBioKG. This two-layer system grounds unseen molecules in biomedical evidence via multi-resolution structural anchoring. It connects an index of 2.74 million molecules (represented by scaffolds, fragments, functional groups, and fingerprints) to a 9.6-million-edge KG. Given only a SMILES string, MolBioKG retrieves structurally related graph entities and traverses their biomedical neighborhoods without task-specific training. It features two inference mechanisms: static multi-anchor retrieval using Reciprocal Rank Fusion, and Adapt-KG, a tool-using LLM policy for adaptive traversal. Evaluated across in-graph link recovery, complex multi-hop reasoning, and out-of-graph generalization, MolBioKG outperforms strong baselines. Notably, it raises Hits@10 from 0.585 to 0.876 in multi-hop reasoning and out-of-graph target recall from 0.145 to 0.269, all while ensuring predictions retain traceable structural anchors and source-attributed KG evidence.