发表机构
Department of Computer Science; University of Milano(计算机科学系; 米兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对生物医学注释验证瓶颈,提出利用生物医学知识图谱估计候选注释合理性的框架。通过基于社区的负采样策略训练分类器,引入合理性度量,实验表明该方法提高了分类器鲁棒性和注释优先级排序有效性,提升了人工智能辅助生物医学编目的效率。
AI 中文摘要
生物医学知识的快速增长使自动生成的生物注释验证成为生物医学编目的主要瓶颈。计算方法能快速生成大量候选注释,但确定其生物学有效性仍需昂贵的专家评审。机器学习技术可利用生物医学知识图谱(bioKGs)支持这一过程。本文提出一个利用bioKGs估计候选注释合理性并指导专家编目的框架。从知识图谱嵌入开始,用基于社区的负采样策略训练特定关系的二元分类器以获得可靠的置信估计。然后引入一系列合理性度量,结合分类器置信度、可靠性及同一对生物实体替代关系提供的语义上下文。实验结果表明,该负采样策略提高了分类器鲁棒性,合理性度量优于单独的分类器置信度,提升了候选注释优先级排序的有效性,表明利用bioKGs提高了人工智能辅助生物医学编目的效率,同时保留专家对最终注释评估的控制权。
英文摘要
The rapid growth of biomedical knowledge has made the validation of automatically generated biological annotations a major bottleneck in biomedical curation. While computational methods can rapidly produce large numbers of candidate annotations, determining which are biologically valid still requires costly expert review. Prioritizing these candidates before manual curation has therefore become a fundamental challenge. Machine learning techniques can support this process by exploiting biomedical knowledge graphs (bioKGs), which capture biological entities and their functional associations. In this work, we propose a framework that leverages bioKGs to estimate the plausibility of candidate annotations and guide expert curation. Starting from knowledge graph embeddings, we train relation-specific binary classifiers using a community-based negative sampling strategy to obtain reliable confidence estimates. We then introduce a family of plausibility measures that combine classifier confidence, classifier reliability, and the semantic context provided by alternative relationships involving the same pair of biological entities. Unlike conventional confidence estimation, the proposed approach explicitly accounts for multiple biologically meaningful relations that may coexist between the same entities. Experimental results on five large bioKGs demonstrate that the proposed negative sampling strategy consistently improves classifier robustness, increasing balanced accuracy by an average of 5.8%. Moreover, the plausibility measures outperform classifier confidence alone, enabling more effective prioritization of candidate annotations for expert review. Overall, our results show that the use of bioKGs improves the efficiency of AI-assisted biomedical curation while preserving expert control over the final annotation assessment.