发表机构
Iowa State University; Brac University(爱荷华州立大学; 布拉克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
BLADE通过蒸馏LLM判断作为冻结教师正则化器,分离潜在真相与图谱记录,在五个基准上显著降低校准误差,同时保持排序竞争力,并仅对声明候选分布保证校准。
AI 中文摘要
知识图谱补全模型优化排序,但许多下游应用需要校准的概率。我们提出BLADE,一种变分模型,将潜在真相与图谱记录分离,并将离线语言模型判断蒸馏到冻结的教师正则化器中。推理期间不涉及LLM。后验样本提供预测概率和认知不确定性,而紧凑的教师仅作为可选分流因素可用。在五个基准上,BLADE在常见排序协议下保持竞争力,并将自适应ECE相对于深度集成平均降低60.1%,相对于温度缩放的RotatE降低78.1%。在相同的FB15k-237候选集上,BLADE在ECE、Brier分数和NLL方面也优于验证选择的直方图分箱和匹配的生成式ComplEx2模型,这些改进在预设的近缺失池中持续存在。在受控注入缺失下,完整分流分数达到平均AUC-PR 0.863,而其最强的非教师变体为0.805。泄漏压力测试表明,对齐语义很重要,但不能排除LLM预训练期间获得的知识。因此,我们仅对声明的候选分布声称校准,而非对所有未观察到的三元组。
英文摘要
Knowledge graph completion models optimize ranking, although many downstream applications require calibrated probabilities. We present BLADE, a variational model that separates latent truth from graph recording and distills offline language-model judgments into a frozen teacher regularizer. The LLM is absent during inference. Posterior samples provide predictive probabilities and epistemic uncertainty, while the compact teacher remains available only as an optional triage factor. Across five benchmarks, BLADE remains competitive under a common ranking protocol and reduces adaptive ECE by a macro-average of 60.1% relative to deep ensembles and 78.1% relative to temperature-scaled RotatE. On identical FB15k-237 candidate sets, BLADE also improves ECE, Brier score, and NLL over validation-selected histogram binning and a matched generative ComplEx2 model, with these improvements persisting on a prespecified near-miss pool. Under controlled injected missingness, the full triage score achieves a mean AUC-PR of 0.863, compared with 0.805 for its strongest non-teacher variant. Leakage stress tests show that aligned semantics matter, but they cannot exclude knowledge acquired during LLM pretraining. We therefore claim calibration only for the declared candidate distributions, not for all unobserved triples.