发表机构
AI Center-Mountain View, Samsung Electronics(三星电子山景城人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对搜索增强型大语言模型智能体结果级奖励监督不足的问题,提出PROGRESS方法,通过教师引导的覆盖奖励结合R1框架训练,提升了智能体任务性能,凸显了搜索策略监督的重要性。
AI 中文摘要
现有搜索增强型大语言模型智能体采用强化学习训练以提升推理能力,但这些方法主要依赖结果级奖励,对搜索行为的监督极少,且忽略智能体正确分解复杂查询的能力。为缓解该问题,我们提出PROGRESS,利用教师引导的覆盖奖励显式塑造策略模型的分解查询生成。训练期间,冻结的教师模型用于将复杂查询分解为必要的搜索查询,这些查询用于引导策略模型的搜索行为。将该方法集成到R1风格训练框架中,可为查询分解决策提供轻量指导,无需密集的过程级监督。实验表明,覆盖引导的强化学习提升了整体任务性能,凸显了在智能体大语言模型中显式监督搜索策略的重要性。
英文摘要
Existing search-augmented LLM agents are trained using Reinforcement Learning to boost its reasoning capabilities. However, these approaches primarily rely on outcome-level rewards, which provide little supervision over search behavior and overlook agent's ability to decompose complex queries properly. To mitigate this issue, we propose PROGRESS which utilizes teacher-guided coverage reward to explicitly shape decomposed query generation of the policy model. During training, frozen teacher models are used to decompose complex queries into essential search queries. These essential search queries are utilized to guide the search behavior of the policy model. Integrated into an R1-style training framework, our approach provides lightweight guidance over query decomposition decisions without dense process-level supervision. Experiments show that coverage-guided RL improves overall task performance, highlighting the importance of explicitly supervising search strategies in agentic LLMs.
CommentsAccepted in Interspeech 2026