arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgenticRag-R1:用于多步推理、检索与记忆的带栈式记忆的智能体强化学习

AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing

Xinke Jiang, Yue Fang, Zhibang Yang, Jiaran Gao, Zhixin Zhang, Tao Feng, Rihong Qiu, Wentao Zhang, Hongxin Ding, Ruizhe Zhang, Yongxin Xu, Yuheng Huang, Xu Chu, Junfeng Zhao, Yasha Wang

arXiv 2608.29622首次发表:更新:

发表机构

National Engineering Research Center of Software Engineering, Peking University; School of Computer Science, Peking University(北京大学国家软件工程研究中心; 北京大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有智能体RAG方法的奖励分配弱、偏向短程推理等问题,提出AgenticRag-R1框架,通过记忆栈等实现长程学习,在多类基准上性能优于强基线。

AI 中文摘要

检索增强生成(Retrieval-Augmented Generation,RAG)提升了大语言模型(Large Language Models,LLMs)的事实性,但现有RAG系统常难以应对需要自适应检索和持续修订中间上下文的复杂多步推理。近期基于强化学习(Reinforcement Learning,RL)的智能体RAG方法部分缓解了该问题,但通常依赖粗粒度动作空间和轨迹级奖励,导致奖励分配薄弱,且存在偏向短 horizon(时间范围)、刻板推理模板的偏差。为解决该问题,我们提出AgenticRag-R1,这是一种通过记忆栈和细粒度动作空间深度整合推理、检索与记忆的RL框架,辅以分层动作感知奖励和信息感知轨迹拒绝策略,以实现有效的长 horizon 学习。在涵盖多种骨干模型规模的多样多跳、开放域及智能体推理基准上的实验表明,AgenticRag-R1始终优于强基线,且能学习到更鲁棒、可解释和记忆感知的推理行为,凸显了细粒度动作建模与信息感知优化对长 horizon 推理的作用。我们的代码可在该https URL匿名获取。

英文摘要

Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs), yet existing RAG systems often struggle with complex, multi-step reasoning that requires adaptive retrieval and continuous revision of intermediate contexts. Recent reinforcement learning (RL)-based agentic RAG methods partially alleviate this issue, but typically rely on coarse-grained action spaces and trajectory-level rewards, resulting in weak reward assignment and a bias toward short-horizon, stereotyped reasoning template. To address, we propose AgenticRag-R1, a RL framework that deeply integrates reasoning, retrieval, and memory via a memory stack and fine-grained action space, supported by hierarchical action-aware rewards and an information-aware trajectory rejection strategy to enable effective long-horizon learning. Experiments across a diverse set of multi-hop, open-domain, and agentic reasoning benchmarks, spanning multiple backbone model sizes, demonstrate that AgenticRag-R1 consistently outperforms strong baselines. Moreover, AgenticRag-R1 learns more robust, interpretable, and memory-aware reasoning behaviors, highlighting the effect of fine-grained action modeling and information-aware optimization for long-horizon reasoning. Our code is anonymous available at https://github.com/jiangxinke/Harness-RL/tree/AgenticRAG-R1-Whitebox.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑