arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BRIDGE:双层检索信用感知的智能体强化学习

BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning

Quan Xiao, Mingda Liu, Gaowen Liu, Katsuki Fujisawa, Tianyi Chen

arXiv 2609.36505首次发表:更新:

发表机构

Cornell University; Institute of Science Tokyo; Cisco Research(康奈尔大学; 东京科学大学; 思科研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出BRIDGE,一种双层优化方法,联合训练检索器与LLM策略以解决智能体强化学习中的信息信用缺口,在七个开放域问答基准上取得最高平均准确率,多跳EM分别提升9.6和3.4点。

AI 中文摘要

具有可验证奖励的智能体强化学习(ARL)通过学会将搜索与推理交错进行,提升了大语言模型(LLM)处理知识密集型任务的能力。然而,大多数现有ARL方法仅优化LLM生成的令牌,并将检索到的证据视为环境观测。这造成了信息信用缺口:由缺失或误导性证据导致的失败被归因于LLM策略而非检索器,这促使我们联合训练LLM和检索器。在本文中,我们表明检索与LLM策略学习是顺序敏感的:在优化策略之前调整检索器比相反顺序产生更大的奖励增益。为了在允许两个组件共同适应的同时保持这种层级结构,我们将检索增强的智能体RL表述为一个双层优化问题。为了高效求解该问题,我们引入了BRIDGE,一种受RL和检索目标损失景观分析启发的内存高效一阶双层方法。在七个开放域问答基准上,BRIDGE在3B和7B骨干网络上均取得了最高平均准确率,分别将多跳平均成绩较最强基线提升了9.6和3.4个EM点。它还在医学问答基准上取得了最佳的平均答案准确率和推理质量。

英文摘要

Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM-generated tokens and treat retrieved evidence as environment observations. This creates an information-credit gap: failures caused by missing or misleading evidence are attributed to the LLM policy rather than to the retriever, which motivates training the LLM and the retriever jointly. In this paper, we show that retrieval and LLM policy learning are order-sensitive: adapting the retriever before optimizing the policy yields a larger reward gain than the reverse order. To preserve this hierarchy while allowing both components to co-adapt, we formulate retrieval-augmented agentic RL as a bilevel optimization problem. To solve it efficiently, we introduce BRIDGE, a memory-efficient first-order bilevel method motivated by a loss-landscape analysis of the RL and retrieval objectives. Across seven open-domain QA benchmarks, BRIDGE achieves the highest average accuracy with both 3B and 7B backbones, improving the multi-hop average over the strongest baseline by 9.6 and 3.4 EM points, respectively. It also achieves the best averaged answer accuracy and reasoning quality across medical QA benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑