arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.31137cs.AI

OntoAligner-Ensemble:基于投票的异构本体对齐技术融合

OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques

Hamed Babaei Giglou, Sören Auer, Peio Popov, Mahsa Sanaei, Jennifer D'Souza

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出OntoAligner-Ensemble框架,通过两阶段投票融合策略整合异构OA对齐器,经OAEI八任务验证,可提升精确率-召回率平衡,为OA提供实用集成选择指导。

中文摘要 AI 辅助

本体对齐(OA)已发展出多种方法范式,从词汇型、结构型对齐器,到知识图谱嵌入(KGE)模型,再到近期基于大语言模型(LLM)的方法。尽管现代OA框架为部署这些异构对齐器提供了统一生态系统,但系统协调其互补且有时相互冲突的预测的机制仍相对未被充分探索。我们提出OntoAligner-Ensemble,这是一个模块化、与对齐器无关的框架,通过可配置的两阶段流程组合候选对应关系:基于投票的融合策略,以及融合后选择策略。该框架支持任何在OntoAligner中实现、能生成候选对应关系的对齐器,使不同对齐范式可通过统一决策流程集成。为验证其有效性,我们使用代表性轻量字符串对齐器、基于KGE的对齐器,以及由开源权重和基于API的LLM驱动的检索增强生成对齐器实例化该框架。我们在五个OAEI赛道的八个基准任务上评估了单个对齐器及集成配置,任务范围从生物医学到非等价对齐。结果表明,集成融合始终能提升精确率与召回率的平衡,且在不同领域中常优于单独对齐器。此外,我们的分析显示,集成组成直接影响精确率-召回率权衡:异构跨范式集成通常提升精确率,而同构LLM集成更常实现更高的整体F1分数。这些发现表明,系统集成学习为OA提供了稳健且可复现的策略,同时为不同对齐场景下的集成组成选择提供了实用指导。

英文摘要

Ontology alignment (OA) has evolved through several methodological paradigms, ranging from lexical and structural aligners to knowledge graph embedding (KGE) models and, more recently, Large Language Model (LLM)-based approaches. Although modern OA frameworks provide unified ecosystems for deploying these heterogeneous aligners, mechanisms for systematically reconciling their complementary and sometimes conflicting predictions remain relatively underexplored. We present OntoAligner-Ensemble, a modular and aligner-agnostic framework that combines candidate correspondences through a configurable two-stage process comprising voting-based fusion strategies followed by post-fusion selection policies. The framework supports any aligner implemented within OntoAligner that produces candidate correspondences, enabling diverse alignment paradigms to be integrated through a unified decision process. To demonstrate its effectiveness, we instantiate the framework using representative lightweight string-aligner, KGE-based, and Retrieval-Augmented Generation aligners powered by both open-weight and API-based LLMs. We evaluate individual aligners and ensemble configurations across eight benchmark tasks from five OAEI tracks spanning biomedical to beyond-equivalence. The results show that ensemble fusion consistently improves the balance between precision and recall and frequently outperforms standalone aligners across diverse domains. Furthermore, our analysis reveals that ensemble composition directly affects the precision-recall trade-off: heterogeneous cross-paradigm ensembles generally improve precision, whereas homogeneous LLM ensembles more often achieve higher overall F1-scores. These findings demonstrate that systematic ensemble learning offers a robust and reproducible strategy for OA while providing practical guidance for selecting ensemble compositions under different alignment scenarios.

发表机构

  • TIB Leibniz Information Centre for Science and Technology(莱布尼茨科学与技术信息中心)
  • L3S Research Center, Leibniz University of Hannover(汉诺威莱布尼茨大学L3S研究中心)
  • Graphwise
  • University of Tabriz(大不里士大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑