arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OPDSearch+:用于搜索增强推理的带RL优化的在线策略蒸馏

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

Qinglin Ye, Zhiyuan Gu, Jingjie Xia, Yiheng Zhang, Kaiyan Zhao, Shunchao Zheng, Yuhang Mu, Wenchao Du, Yiming Wang

arXiv 2608.24310首次发表:更新:

发表机构

University of Chinese Academy of Sciences; Institute of Computing Technology, Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; University of Macau; Wuhan University; Georgia Institute of Technology; Northwestern Polytechnical University; University of Hong Kong(中国科学院大学; 中国科学院计算技术研究所; 中国科学院自动化研究所; 澳门大学; 武汉大学; 佐治亚理工学院; 西北工业大学; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OPDSearch+是无需教师微调的搜索增强推理蒸馏范式,利用冻结的现成指令模型分两阶段优化3B模型,在7个QA基准上优于所有3B RL基线,HotpotQA和2WikiMultihopQA分别提升13.1%、8.5%。

AI 中文摘要

小型语言模型的搜索增强推理仍存在困难。从训练好的教师模型进行在线策略蒸馏(OPD)提供了一个有前景的方向,但存在两个问题:(1)高质量的多轮搜索轨迹依赖于动态检索器响应,使得监督微调(SFT)数据的大规模收集成本过高;(2)针对特定任务训练的教师模型会产生大量训练成本,而直接使用现成教师模型进行OPD且不进行特定任务微调,会将学生模型限制在教师模型的性能上限,且存在严重的训练不稳定性。我们提出OPDSearch+,这是一种无需对搜索增强推理的教师模型进行微调的蒸馏范式。我们研究了在线策略蒸馏中使用冻结的现成指令模型作为教师的作用,并揭示了一个关键见解:教师会重塑学生的策略分布,使得后续的强化学习(RL)能够收敛到仅靠RL无法达到的更优解决方案。在第一阶段,学生与实时搜索引擎交互,并通过逐位置前向KL目标进行蒸馏,无需任何特定任务的教师训练即可传递推理分解和证据整合技能。在第二阶段,RL从更丰富的行为基础上优化蒸馏后的学生,达到仅靠RL从头开始无法实现的性能。在七个问答基准测试中,采用3B模型的OPDSearch+始终优于所有先前的3B RL基线,在HotpotQA上取得13.1%的提升,在2WikiMultihopQA上取得8.5%的提升。

英文摘要

Search-augmented reasoning remains difficult for small language models. On-policy distillation (OPD) from trained teachers offers a promising direction, but suffers from two issues: (1) high-quality multi-turn search trajectories depend on dynamic retriever responses, making SFT data prohibitively expensive to collect at scale; (2) task-specifically trained teachers incur substantial training cost, while directly applying OPD with an off-the-shelf teacher without task-specific fine-tuning constrains the student to the teacher's performance ceiling and suffers from severe training instability. We propose OPDSearch+, the first distillation paradigm that requires no teacher fine-tuning for search-augmented reasoning. We investigate the role of a frozen off-the-shelf instruct model as the teacher in on-policy distillation, and reveal a key insight: the teacher reshapes the student's policy distribution so that subsequent RL converges to a superior solution that RL alone cannot reach. In stage one, the student interacts with a live search engine and is distilled via a per-position forward KL objective, transferring reasoning decomposition and evidence integration skills without any task-specific teacher training. In stage two, RL refines the distilled student from a richer behavioral foundation, achieving performance that RL alone cannot reach from scratch. Across seven QA benchmarks, OPDSearch+ with a 3B model consistently outperforms all prior 3B RL baselines, achieving gains of 13.1% on HotpotQA and 8.5% on 2WikiMultihopQA.

Comments9 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑