arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29518cs.LGcs.CE

CataOPD:用于大语言模型推理的催化式在策略蒸馏

CataOPD: Catalytic On-Policy Distillation for Large Language Model Reasoning

  • Nanyang Technological University(南洋理工大学)
  • Hithink Research(海天瑞声研究院)
  • Nanjing University of Science and Technology(南京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Wenjin Liu, Chenxi Wang, Jiapu Wang, Zhe Cui, Anh Tuan Luu, Haoran Luo

AI总结:

CataOPD通过教师作为催化剂而非目标,结合自救援路由、催化引导自求解和屏障加权内化,扩展学生推理可达性,提升大语言模型推理性能与泛化能力。

AI中文摘要:

强化学习(RL)和在策略蒸馏(OPD)是提升大语言模型推理能力的两种代表性范式。然而,当未采样到正确轨迹时,RL缺乏正面的正确性信号,而OPD仍受限于学生在在策略分布下可达到的推理轨迹。因此,我们提出CataOPD,其中教师扮演催化剂而非目标的角色,在扩展可达性的同时,将经过验证的学生生成轨迹内化到无催化剂策略中。自救援路由(Self-Rescue Routing)使用经验上全失败的组作为路由信号而非教师干预触发器,首先通过额外的在策略自采样寻求正确轨迹。对于自救援后仍未解决的问题,催化引导自求解(Catalytic-Guided Self-Resolution)利用催化引导在引导的学生分布中引出经过验证的学生生成轨迹。屏障加权内化(Barrier-Weighted Internalization)根据引导与未引导的对数概率差距对词元进行加权,将更新聚焦于在无引导情况下难以处理的决定性词元。实验结果表明,CataOPD优于现有基线,将独立学生推理扩展到仍未恢复的问题,并在无催化剂推理下改善了分布外泛化。我们的项目可在以下网址获取:https URL。

英文摘要:

Reinforcement learning (RL) and on-policy distillation (OPD) are two representative paradigms for improving large language model reasoning. However, when no correct trajectory is sampled, RL lacks a positive correctness signal, while OPD remains constrained by the reasoning trajectories reachable under the student's on-policy distribution. Therefore, we propose CataOPD, where the teacher acts as a catalyst rather than a target, expanding reachability while internalizing verified student-produced trajectories into a catalyst-free policy. Self-Rescue Routing uses empirically all-failed groups as routing signals rather than teacher-intervention triggers, first seeking correct trajectories through additional on-policy self-sampling. For problems unresolved after self-rescue, Catalytic-Guided Self-Resolution uses catalytic guidance to elicit a verified student-produced trajectory in the guided student distribution. Barrier-Weighted Internalization weights tokens by guided-to-unguided log-probability gaps, focusing updates on decisive tokens difficult without guidance. Experimental results show that CataOPD outperforms current baselines, extends independent student reasoning to still-unrecovered problems, and improves out-of-distribution generalization under catalyst-free inference. Our project is available at https://github.com/QwenQKing/CataOPD.

↑