发表机构
Shenzhen University; Great Bay University; Ant Group(深圳大学; 大湾区大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GTA-RAG是一种图轨迹增强强化学习框架,通过轨迹级监督优化检索策略,在多跳和简单问答基准上,相比RL-based RAG基线提升了性能与证据链覆盖率。
AI 中文摘要
检索增强生成(RAG)使大语言模型(LLM)能够访问外部知识以回答知识密集型问题。对于复杂的多跳问题,多轮检索增强推理将RAG扩展为迭代过程,该过程会反复搜索并整合跨文档的证据。然而,现有的用于智能体RAG的强化学习(RL)方法通常采用最终答案奖励进行优化,这种奖励提供的监督信号稀疏,且忽略了模型是否实际检索到所需的证据链。我们提出GTA-RAG,一种用于多轮检索增强推理的图轨迹增强RL框架。我们从实体-文档图中采样连接的文档路径,合成多跳问答轨迹,并使用部署的检索器验证这些轨迹,以获得可执行的轨迹级监督信号。随后,我们使用组相对策略优化(GRPO)和轨迹引导奖励来优化检索策略,该奖励同时鼓励准确回答和获取目标证据文档,之后在自然问答实例上进行答案奖励训练。在3个多跳和2个简单问答基准上的实验表明,使用Qwen2.5-3B和Qwen2.5-7B作为主干模型时,我们的方法始终优于基于RL的RAG基线,同时大幅提高了证据链覆盖率。我们的代码可在该https URL获取。
英文摘要
Retrieval-augmented generation (RAG) enables LLMs to access external knowledge for answering knowledge-intensive questions. For complex multi-hop questions, multi-turn retrieval-augmented reasoning extends RAG into an iterative process that repeatedly searches for and integrates evidence across documents. However, existing reinforcement-learning (RL) approaches for agentic RAG are typically optimized with final-answer rewards, which provide sparse supervision and overlook whether the model actually retrieves the required evidence chain. We present \textsc{GTA-RAG}, a graph-trajectory-augmented RL framework for multi-turn retrieval-augmented reasoning. From an entity--document graph, we sample connected document paths, synthesize multi-hop QA trajectories, and validate them with the deployed retriever to obtain executable trajectory-level supervision. We then optimize the retrieval policy with Group Relative Policy Optimization (GRPO) and a trajectory-guided reward that encourages both accurate answers and acquisition of target evidence documents, followed by answer-reward training on natural QA instances. Experiments on three multi-hop and two simple QA benchmarks show that \method{} consistently outperforms RL-based RAG baselines with both Qwen2.5-3B and Qwen2.5-7B backbones, while substantially improving evidence-chain coverage. Our code is available at https://github.com/cjcj46262/GTA-RAG.
Comments12 pages, 5 figures. Accepted to EMNLP 2026 Findings