arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12764cs.LGcs.AI

超越结果奖励:面向深度搜索智能体的步骤级自蒸馏策略优化

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

发表机构华为技术有限公司 · 香港科技大学
查看机构详情
  • Huawei Technologies Ltd.(华为技术有限公司)
  • The Hong Kong University of Science Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Haoze Wu, Chuqiao Kuang, Tianyi Zhuang, Xiaoguang Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对深度搜索智能体信用分配稀疏问题,提出 SSPO 方法,通过证据锚点与步骤级优势权重优化策略,在多个数据集上性能优于 GRPO 且开销低。

中文摘要 AI 辅助

深度搜索智能体在数十步的轨迹上运行,但标准强化学习仅为每条轨迹提供单一的结果奖励,对于有效的信用分配而言这过于稀疏。在线策略自蒸馏(OPSD)通过使用模型自身的逻辑作为密集的 token 级教师来解决这一问题,但将其扩展到搜索智能体时会引入一个根本矛盾:教师拥有正确答案等特权信息,其生成的分布与学生基于探索的推理存在系统性差异,而朴素蒸馏会导致学生继承这种信息不对称,而非学习更好的搜索策略。我们通过两项贡献解决这一矛盾:首先,我们构建证据锚点(Evidence Anchors),即从网络中提取的简洁步骤级证据片段,作为捕获关键推理步骤而不泄露整个答案路径的特权信息;其次,我们提出步骤级自蒸馏策略优化(SSPO),该方法将教师-学生分歧转换为 GRPO 内的步骤级优势权重,仅应用于错误轨迹。此设计将更新内容与更新幅度解耦:结果奖励决定策略变化的方向,而教师在每一步调节其幅度。正确轨迹保持不变,以保留其多样性。在 Qwen3-8B 上,SSPO 在 BrowseComp、GAIA 和 FRAMES 数据集上始终优于 GRPO,超过或匹配训练时梯度步数为其两倍的 GRPO,同时每步仅因单次额外前向传播增加约 5% 的开销。

英文摘要

Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory, which is far too sparse for effective credit assignment. On-policy self-distillation (OPSD) addresses this by using the model's own logits as dense token-level teachers, but extending it to search agents introduces a fundamental tension: the teacher, having access to privileged information such as the correct answer, produces a distribution that differs systematically from the student's exploration-based reasoning, and naive distillation causes the student to inherit this information asymmetry rather than learn better search strategies. We resolve this tension through two contributions. First, we construct Evidence Anchors, which are concise, step-level evidence snippets extracted from the web, as privileged information that captures key reasoning steps without revealing the entire answer path. Second, we propose Step-Level Self-Distilled Policy Optimization (SSPO), which converts teacher-student disagreement into step-level advantage weights within GRPO, applied exclusively to incorrect trajectories. This design decouples what to update from how much to update: the outcome reward determines the direction of policy change, while the teacher modulates its magnitude at each step. Correct trajectories are left untouched, preserving their diversity. On Qwen3-8B, SSPO consistently outperforms GRPO across BrowseComp, GAIA, and FRAMES, surpassing or matching GRPO trained with twice as many gradient steps while adding only about 5 percent overhead per step from a single additional forward pass.

补充信息

↑