arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OntologyBench:稠密检索能否满足结构化生物医学约束?

OntologyBench: Can Dense Retrieval Satisfy Structured Biomedical Constraints?

Xiao Yu Cindy Zhang, Wyeth Wasserman, Jian Zhu

arXiv 2609.08174首次发表:更新:

发表机构

The University of British Columbia; BC Children’s Hospital Research Institute(不列颠哥伦比亚大学; 不列颠哥伦比亚省儿童医院研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出生物医学检索基准OntologyBench,发现稠密检索在关系及组合任务上性能不足,微调有助提升,而重排序和LLM方法改进有限,需结合结构化知识。

AI 中文摘要

我们引入了OntologyBench,一个分层生物医学检索基准,包含471,854个训练和125,744个评估的查询-文档相关性对,涵盖概念接地、关系检索和基于组合表型的检索。尽管这些任务使用具有本体感知的参考方法可以处理,但在各任务层级中,嵌入性能在关系型和组合型任务上通常低于概念接地任务。基于本体派生的监督进行微调可提高若干关系型和组合型任务的性能,而评估的重排序和基于LLM的候选评分方法几乎没有或没有端到端的改进。错误通常反映疾病仅匹配表型证据的子集。这些发现表明,所评估的嵌入和重排序配置无法可靠地恢复所选本体关系和表型组合所编码的兼容性,并促使检索系统更好地整合学习表示与结构化生物医学知识。

英文摘要

We introduce OntologyBench, a tiered biomedical retrieval benchmark comprising 471,854 training and 125,744 evaluation query-document relevance pairs across concept grounding, relational retrieval, and compositional phenotype-based retrieval. Although these tasks can be tractable using ontology-aware reference methods, across task tiers, embedding performance is generally lower on relational and compositional tasks than on concept-grounding tasks. Fine-tuning on ontology-derived supervision improves performance on several relational and compositional tasks, whereas the evaluated reranking and LLM-based candidate-scoring methods provide little or no end-to-end improvement. Errors frequently reflect diseases matching only subsets of the phenotype evidence. These findings indicate that the evaluated embedding and reranking configurations do not reliably recover the compatibility encoded by the selected ontology relations and phenotype combinations and motivate retrieval systems that better integrate learned representations with structured biomedical knowledge.

Comments19 pages and 1 figure. Accepted to EMNLP 2026 Main conference. https://openreview.net/forum?id=RMiOTIub1j&noteId=RMiOTIub1j

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑