知识图谱表示学习中的不确定性
Uncertainty in Representation Learning on Knowledge Graphs
浏览论文内容
中文总结 AI 辅助
该论文系统研究知识图谱嵌入中的三类不确定性,提出投票聚合、共形预测及统计有效区间等方法,实现可靠且不确定性感知的知识图谱推理。
中文摘要 AI 辅助
知识图谱嵌入(KGE)方法将实体和谓词表示在连续向量空间中,以推断缺失的知识。尽管在基准测试中表现强劲,其预测往往缺乏有原则的可靠性保证,限制了其在高风险应用中的使用。此外,不确定性贯穿KGE流程的始终,从输入知识的不完整或概率性到随机训练和预测。本论文系统地研究了KGE中不确定性的三个来源:知识不确定性,源于输入知识的不完整、噪声或概率性;算法不确定性,由模型训练中的随机性引起;以及预测不确定性,涉及模型输出的可靠性。为解决算法不确定性,论文表明在相同设置下训练的模型可能产生显著不同的预测,并引入了一种基于投票的聚合框架以缓解这种不稳定性。为量化预测不确定性,论文将共形预测适配到KGE,构建具有无分布覆盖保证的答案集,并将其扩展以提供谓词条件可靠性保证。为支持知识不确定性下的推理,论文为置信度评分三元组开发了统计上有效的预测区间,以及一种基于嵌入的方法来近似统计本体上的概率推理,并具有形式上的健全性保证。总之,这些互补的、模型无关的方法为KGE中的不确定性提供了一种实用且理论扎实的途径,超越了预测准确性,迈向可靠且具有不确定性意识的知识图谱推理。
英文摘要
Knowledge graph embedding (KGE) methods represent entities and predicates in continuous vector spaces to infer missing knowledge. Despite strong benchmark performance, their predictions often lack principled reliability guarantees, limiting their use in high-stakes applications. Moreover, uncertainty arises throughout the KGE pipeline, from incomplete or probabilistic input knowledge to stochastic training and prediction. This thesis systematically investigates three sources of uncertainty in KGE: knowledge uncertainty, arising from incomplete, noisy, or probabilistic input knowledge; algorithmic uncertainty, induced by randomness in model training; and predictive uncertainty, concerning the reliability of model outputs. To address algorithmic uncertainty, the thesis demonstrates that models trained under identical settings can produce substantially different predictions and introduces a voting-based aggregation framework to mitigate this instability. To quantify predictive uncertainty, it adapts conformal prediction to KGE, constructing answer sets with distribution-free coverage guarantees and extending them to provide predicate-conditional reliability guarantees. To support reasoning under knowledge uncertainty, it develops statistically valid prediction intervals for confidence-scored triples and an embedding-based approach to approximate probabilistic reasoning over statistical ontologies with formal soundness guarantees. Together, these complementary, model-agnostic methods provide a practical and theoretically grounded approach to uncertainty in KGE, advancing beyond predictive accuracy toward reliable and uncertainty-aware knowledge graph reasoning.