Oracle,我会学会吗?链接预测模型间的预测收敛性与互补性研究
Oracle, will I ever learn? A study of prediction convergence and complementarity across link prediction models
另 1 家 · 查看机构详情
- Université Côte d’Azur(蔚蓝海岸大学)
- Inria(法国国家信息与自动化研究所)
- CNRS(法国国家科学研究中心)
- I3S(信息科学与信号处理实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究探究不同链接预测模型的预测差异,通过Oracle衡量模型互补性,发现模型间存在互补性但随模型数量增加会饱和,为Web应用的链接预测模型组合提供了性能上限与局限的结论。
中文摘要 AI 辅助
知识图谱已成为Web应用(包括搜索、问答和推荐系统)中重要的结构化知识来源。在这些应用中,链接预测既可以作为一项预测任务本身,也可以作为一种为下游任务补充不完整知识图谱的手段。有趣的是,不同的链接预测模型,甚至同一模型的不同训练运行,会对同一查询产生截然不同的预测结果。这表明模型在捕获底层知识时存在差异,从而引发一个基本问题:不同模型在多大程度上捕获了互补知识,以及通过组合这些模型可以恢复多少此类知识?我们提议通过一种“Oracle(最优选择器)”的性能来衡量模型互补性,该Oracle会为每个查询从所考虑的一组模型中选择最佳预测,从而提供模型组合可实现性能的上限。在多种架构和基准测试中,我们发现单个模型与它们的Oracle之间存在显著差距,这表明不同模型确实捕获了互补知识。然而,随着添加更多模型,这种互补性会迅速饱和,即使使用大量模型,仍有一部分查询无法被解决。这些发现既揭示了模型互补性的潜力,也揭示了当前链接预测模型集体可恢复知识的基本局限;因此强调需要进一步研究以构建稳健的Web应用。
英文摘要
Knowledge graphs have become an important source of structured knowledge for Web applications, including search, question answering, and recommender systems. In these applications, link prediction can serve either as a prediction task itself or as a means to enrich incomplete knowledge graphs for downstream tasks. Interestingly, different link prediction models, or even different training runs of the same model, can produce substantially different predictions for the same query. This suggests a variability in the capture of the underlying knowledge by models, thus raising a fundamental question: to what extent do different models capture complementary knowledge, and how much of this knowledge could be recovered by combining them? We propose to measure model complementarity through the performance of an oracle that, for each query, selects the best prediction among a considered set of models, hence providing an upper bound on the performance achievable through model combination. Across several architectures and benchmarks, we find a substantial gap between individual models and their oracle, revealing that different models capture complementary knowledge. Yet, this complementarity rapidly saturates as more models are added, leaving a persistent subset of queries unsolved even by a large number of models. These findings reveal both the potential of model complementarity and a fundamental limit to what current link prediction models can collectively recover; thereby highlighting the need for further research to build robust Web applications.