发表机构
Artificial Intelligence Research Center, Chang Gung University; Graduate Institute of Oral Biology, National Taiwan University; National Taiwan University of Science and Technology(长庚大学人工智能研究中心; 国立台湾大学口腔生物学研究所; 国立台湾科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对时态知识图谱推理中外推设置下RL训练奖励稀疏、探索效率低的问题,提出RAPTOR自监督预训练方法,通过学习估计可达性减少无前景路径探索,为RL微调提供初始化,实验证明该方法能提高训练效率并优于传统基线。
AI 中文摘要
时态知识图谱(TKG)推理的外推设置专注于从时态知识图谱中的历史数据预测未来时间戳事件(事实)。现有基于强化学习(RL)的多跳推理方法在TKG推理中很突出,因其通过显式多跳路径追踪产生可人工解释的预测。但RL训练中奖励稀疏,因庞大且随时间演变的动作空间探索效率极低。为应对这些挑战,我们提出RAPTOR,一种自监督预训练方法,向智能体注入可达性感知归纳偏差。通过学习估计候选动作到目标实体的可达性,减少对无前景路径的探索,并为下游RL微调提供强初始化。在ICEWS14、ICEWS05 - 15和ICEWS18数据集上的实验结果表明,RAPTOR预训练显著提高训练效率,持续优于传统基线,是增强基于RL的TKG推理多跳推理方法的有效途径。
英文摘要
Temporal Knowledge Graph (TKG) reasoning under the extrapolation setting focuses on forecasting future time-stamped events (facts) from historical data in a temporal knowledge graph. Existing approaches, reinforcement learning (RL)-based multi-hop reasoning methods are prominent for TKG reasoning because they produce human-interpretable predictions via explicit multi-hop path tracing. However, during RL training, rewards are typically sparse, and exploration is highly inefficient due to the vast, time-evolving action space. These issues hinder efficient training and often limit overall performance. To address these challenges, we propose RAPTOR (Reachability-Aware Pretraining for Efficient Target-Oriented Path Exploration), a self-supervised pretraining method that injects a reachability-aware inductive bias to the agent. By learning to estimate the reachability of candidate actions to the target entity, RAPTOR reduces exploration over unpromising paths and provides a strong initialization for downstream RL fine-tuning. Experimental results on the ICEWS14, ICEWS05-15, and ICEWS18 datasets demonstrate that RAPTOR pretraining markedly improves the training efficiency and consistently outperforms conventional baselines, establishing it as an effective approach for enhancing RL-based multi-hop reasoning methods for TKG reasoning.