发表机构
University of Edinburgh; Huawei Technologies Research & Development (UK) Limited(爱丁堡大学; 华为技术研发(英国)有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对搜索智能体强化学习中稀疏结果奖励导致的信用分配难题,系统研究中间监督信号,提出结合中间信号与结果奖励的训练框架,并在多基准上验证了性能提升,凸显奖励设计与信用分配的关键作用。
AI 中文摘要
搜索智能体使大型语言模型(LLMs)能够迭代检索并利用信息来回答复杂的多跳问题。基于可验证奖励的强化学习(RLVR)为这类智能体的后训练提供了一种有前景的方法,但其依赖于稀疏的、基于结果的监督,这可能导致信用分配困难并限制学习效率。在本文中,我们系统地研究了中间监督如何改进搜索智能体的强化学习。我们研究了一系列从中间检索步骤提供学习信号的奖励塑形和信用分配策略。基于这些见解,我们开发了一个训练框架,将中间信号与最终结果奖励相结合,以改进从多步搜索轨迹中的学习。在匹配的训练条件下,跨多个基准的实验表明,搜索智能体的整体性能得到提升,并显示中间信号的选择及其信用分配的位置都会影响训练行为。这些发现表明,奖励设计和信用分配是训练有效搜索智能体的重要设计维度。
英文摘要
Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.