发表机构
Nankai University; Peking University; University of Electronic Science and Technology of China; Shanghai University of Finance and Economics(南开大学; 北京大学; 电子科技大学; 上海财经大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究时态知识图谱问题生成,提出ChronoQG框架,通过整合多种方法构建有时态表现力和跳数限制的基准数据集,评估多种设置下的相关方法,揭示静态KGQG与TKGQG差距,为时态忠实问题生成提供测试平台。
AI 中文摘要
知识图谱问题生成(KGQG)旨在从结构化图谱证据生成自然语言问题。然而,现有KGQG基准大多基于静态知识图谱构建,未编码图谱事实的时态范围。因此无法评估生成问题是否忠实保持时态有效性、事件顺序和答案确定的时态约束。本文研究时态知识图谱问题生成(TKGQG),提出ChronoQG,首个用于TKGQG的有时态表现力和跳数限制的基准构建框架。ChronoQG整合全面的时态约束分类法、拓扑 - 时态子图采样和基于轨迹的问题生成来构建时态忠实问题。该框架从异构时态知识图谱生成四个基准数据集,共16011个经验证问题。评估多种TKGQG设置下基于大语言模型的KGQG方法和提示基线,结果表明现有方法难以保持时态约束。这些发现揭示了静态KGQG和TKGQG之间的明显差距,并确立ChronoQG为具有挑战性的时态忠实问题生成测试平台。
英文摘要
Knowledge graph question generation (KGQG) aims to generate natural-language questions from structured graph evidence. Existing KGQG benchmarks, however, are mostly built on static knowledge graphs and do not encode the temporal scopes of graph facts. As a result, they cannot evaluate whether generated questions faithfully preserve temporal validity, event ordering, and answer-determining temporal constraints. In this paper, we study temporal knowledge graph question generation (TKGQG), where a generated question must be faithful to both the support subgraph and the temporal constraints required to identify the target answer. We propose ChronoQG, the first temporally expressive and hop-bounded benchmark construction framework for TKGQG. ChronoQG integrates a comprehensive temporal-constraint taxonomy, topology-temporal subgraph sampling, and trace-grounded question generation to construct temporally faithful questions. The framework produces four benchmark datasets from heterogeneous temporal knowledge graphs, totaling 16,011 verified questions. We evaluate representative LLM-based KGQG methods and prompting baselines across diverse TKGQG settings, including temporal-constraint counts, topological templates, and temporal-constraint types. The results show that existing methods struggle to preserve temporal constraints, especially under multi-constraint settings and harder temporal-constraint types. These findings reveal a clear gap between static KGQG and TKGQG, and establish ChronoQG as a challenging testbed for temporally faithful question generation.
CommentsPreprint