arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23214cs.CLcs.IR

生物医学文本与知识图谱的对齐:轻量对齐策略的系统比较

Aligning Biomedical Texts and Knowledge Graphs: A Systematic Comparison of Lightweight Alignment Strategies

Artem Bisliouk, Elizaveta Nosova, Heiko Paulheim, Andreea Iana, Rita T. Sousa

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出统一框架系统比较生物医学文本与KG的轻量对齐策略,构建CTD-Align语料库,发现三元组组合和训练方向影响最大,线性投影效果最优,为二者桥接提供实用基础。

中文摘要 AI 辅助

生物医学知识以两种互补但截然不同的形式存在:非结构化的科学文献与结构化的知识图谱(KG)。对齐二者对知识 grounding、证据检索和KG补全至关重要,但现有方法未明确将自由文本证据与KG三元组对齐。我们提出一个统一框架,系统研究生物医学文本与KG对齐的设计选择。在文本编码器和KG嵌入模型均冻结的情况下,我们仅通过对比学习目标学习二者空间间的轻量投影,这可对六个设计维度进行公平比较:文本编码器、KG嵌入模型、投影头、三元组组合、训练方向和难负样本采样。我们构建了CTD-Align语料库,包含超过2.2万对一一对应的三元组-文献对,将比较毒物基因组数据库(Comparative Toxicogenomics Database)中的化学-基因相互作用与支持性PubMed段落关联。我们在两种检索设置下对其进行评估:文献到三元组和三元组到文献。我们发现三元组组合和训练方向(即共享检索空间)影响最大,而文本编码器和难负样本采样影响很小。总体而言,简单选择效果最佳:将文本投影到KG空间,采用线性头连接主语、谓词和宾语嵌入的表现最优。这些发现确立了轻量对比对齐作为桥接生物医学文本与KG的有效、实用基础。

英文摘要

Biomedical knowledge exists in two complementary but distinct forms: unstructured scientific literature and structured knowledge graphs (KGs). Aligning them is essential for knowledge grounding, evidence retrieval, and KG completion, yet existing methods do not explicitly align free-text evidence with KG triples. We present a unified framework for systematically studying design choices for aligning biomedical text and KGs. With a text encoder and a KG embedding model both frozen, we learn only a lightweight projection between their spaces via a contrastive objective. This enables a fair comparison across six design dimensions: text encoder, KG embedding model, projection head, triple composition, training direction, and hard-negatives sampling. We construct CTD-Align, a corpus of over 22K one-to-one tripledocument pairs linking chemical-gene interactions from the Comparative Toxicogenomics Database to supporting PubMed passages. We evaluate alignment on it in two retrieval settings: document-to-triple and triple-to-document. We find that the triple composition and the training direction (i.e., shared retrieval space) have the greatest impact, whereas the text encoder and hard-negatives sampling matter little. Overall, simple choices win: projecting text into the KG space with a linear head over concatenated subject, predicate, and object embeddings performs best. These findings establish lightweight contrastive alignment as an effective, practical foundation for bridging biomedical text and KGs.

发表机构

  • University of Mannheim(曼海姆大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑