arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

搜索图-R1:使用强化学习训练大型语言模型以搜索知识图谱

Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning

Jia Ao Sun, Hao Yu, Fengran Mo, Zhan Su, Yuchen Hui, Bang Liu, Jian-Yun Nie

arXiv 2607.18481首次发表:更新:

发表机构

Université de Montréal; Mila – Québec AI Institute; McGill University; Halmstad University College(蒙特利尔大学; 米拉-魁北克人工智能研究所; 麦吉尔大学; 哈尔姆斯塔德大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对知识图谱问答部署成本高的问题,提出搜索图-R1,通过监督微调与强化学习,将导航内化到8B模型。核心是用黄金SPARQL查询搭建教师遍历答案路径。该模型在多个数据集上超越前沿语言模型,推理无需辅助模块,训练无需语言模型评判,且训练阶段互补,可跨模型家族迁移。

AI 中文摘要

知识图谱问答(KGQA)需要从主题实体出发,通过多个关系找到答案。近期方法借助检索工具促使前沿语言模型探索图谱,但依赖前沿规模推理使其部署成本高昂。我们提出搜索图-R1(\sogrone{}),通过监督微调(SFT)然后强化学习(RL),将这种导航内化到一个紧凑的8B模型中。核心思想是用每个问题的黄金SPARQL查询搭建前沿教师,教师用实时的\texttt{Search}工具遍历已知的带答案路径。在WebQSP、CWQ和GrailQA数据集上,8B的\sogrone{}超越了所有对比的冻结前沿语言模型系统,在CWQ上取得最强结果,推理时无需辅助模块,训练时无需语言模型评判。隔离每个训练阶段表明SFT和RL有互补收益,方法可跨模型家族迁移,RL比SFT初始化时能用更少的\texttt{Search}调用找到答案。

英文摘要

Knowledge graph question answering (KGQA) requires navigating from topic entities to an answer several relations away. Recent methods prompt a frontier LLM to explore the graph through a retrieval tool, but their reliance on frontier-scale inference makes them costly to deploy. We present Search-on-Graph-R1 (\sogrone{}), which internalizes this navigation into a compact 8B model through supervised fine-tuning (SFT) followed by reinforcement learning (RL). Our central idea is to scaffold a frontier teacher with each question's gold SPARQL query, so the teacher traverses a known answer-bearing path with a live \texttt{Search} tool rather than having to discover the path itself. Since every call executes against a live Freebase server, the resulting trajectories are grounded in the knowledge graph by construction. On WebQSP, CWQ, and GrailQA, \sogrone{} at 8B surpasses every frozen frontier-LLM system in our comparison and posts the strongest results on CWQ of any system we compare against. It does so using no auxiliary module at inference and no LLM judge during training. Isolating each training stage shows that SFT and RL contribute complementary gains, our approach transfers across model families, and RL learns to reach answers in fewer \texttt{Search} calls than its SFT initialization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑