先修复再增强:上下文增强的知识图谱推理用于多跳问答
Repair Before Reinforce: Context-Augmented Knowledge Graph Reasoning for Multi-Hop Question Answering
- Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出上下文增强训练框架,通过附加支持三元组形成上下文图,结合自适应修复和强化学习,提升多跳问答性能。
AI中文摘要:
问答通常需要跨多个相连事实进行推理,而非检索单个孤立关系。知识图谱(KGs)提供了一种结构化表示此类事实的方式,但仅使用孤立的KG头-关系-尾三元组训练大型语言模型(LLMs)可能限制其学习多跳推理所需的周围上下文的能力。在这项工作中,我们提出了一种用于多跳问答的上下文增强训练框架。尽管该框架具有普遍适用性,我们在针对胃轻瘫和糖尿病的疾病特定知识图谱背景下验证了该框架,这些知识图谱是使用一个名为GraphMERT的可靠KG提取框架提取的。对于每个主KG三元组,我们附加从同一源文本块中提取的支持三元组,以形成上下文图(CG)。这创建了两种监督设置:KG基础监督,仅使用目标KG三元组或路径;以及CG基础监督,使用目标KG三元组或路径连同支持上下文三元组。我们在两种设置下使用监督微调(SFT)训练Qwen3-14B模型,生成KGModel和CGModel变体。为了加强模型的较低跳事实基础,我们引入了一个由LLM评判的、历史感知的自适应修复流程,该流程识别未解决的一跳失败,持续对有针对性的修复示例进行微调,并移除或隔离有问题的噪声三元组。这一修复阶段使模型在清理后的保留一跳验证集上达到100%的准确率。最后,我们使用较低跳的问答项进行强化学习(RL),并评估在更难的3跳、4跳和5跳任务上的泛化能力。在两种疾病中,上下文增强监督持续优于仅KG监督的多跳性能。从修复后的SFT检查点初始化的RL产生了更大且更稳定的收益。
英文摘要:
Question-answering often requires reasoning across multiple connected facts rather than retrieving a single isolated relation. Knowledge graphs (KGs) provide a structured way to represent such facts, but training large language models (LLMs) only on isolated KG head-relation-tail triples may limit their ability to learn the surrounding context needed for multi-hop reasoning. In this work, we propose a context-augmented training framework for multi-hop question-answering. Although generally applicable, we validate the framework in the context of disease-specific KGs, extracted using a reliable KG extraction framework called GraphMERT, for Gastroparesis and Diabetes. For each primary KG triple, we attach supporting triples extracted from the same source text chunk to form a context graph (CG). This creates two supervision settings: KG-grounded supervision, which uses only the target KG triple or path, and CG-grounded supervision, which uses the target KG triple or path together with supporting context triples. We train the Qwen3-14B model using supervised fine-tuning (SFT) under both settings, producing KGModel and CGModel variants. To strengthen the lower-hop factual foundation of the models, we introduce an LLM-judged, history-aware adaptive repair pipeline that identifies unresolved one-hop failures, continually fine-tunes on targeted repair examples, and removes or quarantines problematic noisy triples. This repair stage enables the models to reach 100% accuracy on the cleaned retained one-hop validation sets. Finally, we employ reinforcement learning (RL) using lower-hop question-answer items and evaluate generalization on harder 3-hop, 4-hop, and 5-hop tasks. Across both diseases, context-augmented supervision consistently improves multi-hop performance over KG-only supervision. RL initialized from repaired SFT checkpoints yields larger and more stable gains.