AI 中文总结
本研究针对知识图谱学习中数据与理论问题,提出含无监督预训练与监督学习的两阶段框架,建立非渐近风险界,经合成与真实实验验证其有效性。
AI 中文摘要
知识图谱学习是表示与推理结构化知识的强大框架,具有广泛实际应用,但每个实体对应的特定关系标注三元组稀缺,阻碍了高表达能力模型的训练;评分函数的特设设计则限制了模型的泛化性,且缺乏理论基础。我们针对这两个问题,提出了一个具有理论基础的端到端训练框架,该框架扩展并包含现有方法。我们的框架是两阶段流程:在异构语料库上进行无监督预训练,随后针对多种关系类型开展监督学习。我们建立了非渐近风险界,将预训练表示误差与标注样本复杂度解耦,从理论上量化了大规模未标注数据对下游知识预测的益处。合成实验验证了各理论组成部分,真实世界实验也证实了我们的方法在大规模知识图谱基准上的有效性。
英文摘要
Knowledge graph learning provides a powerful framework for representing and inferring structured knowledge, with broad practical applications. However, the scarcity of relation-specific labeled triples per entity hinders the training of expressive models, and the ad hoc design of scoring functions limits generalizability and lacks theoretical grounding. We address both issues with a theoretically grounded, end-to-end training framework that extends and subsumes existing methods. Our framework is a two-stage procedure: unsupervised pretraining over heterogeneous corpora followed by supervised learning with multiple relation types. We establish a nonasymptotic risk bound that disentangles pretraining representation error from labeled-sample complexity, formally quantifying the benefit of large-scale unlabeled data for downstream knowledge prediction. Synthetic experiments validate each theoretical component, and real-world experiments confirm the effectiveness of our approach on large-scale knowledge graph benchmarks.
Comments49 pages, 5 figures, 7 tables