arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

代理探索与可重用引导:基于代理引导更新信号的模块化大语言模型训练后范式

Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update

Daocheng Fu, Rong Wu, Yu Yang, Jianbiao Mei, Licheng Wen, Pinlong Cai, Xuemeng Yang, Yong Liu, Botian Shi, Yu Qiao

arXiv 2607.11505首次发表:更新:

AI 中文总结

研究针对大语言模型训练后优化信号耦合问题,提出PUST框架,通过代理模型探索、提取并传输相对改进信号,解耦更新信号探索与分布对齐,减少计算开销,支持异步处理与跨模型转移,提升训练后范式的模块化、可重用性与效率。

AI 中文摘要

训练后对于提升大语言模型(LLMs)特定领域能力至关重要,但现有奖励优化和分布匹配方法将策略探索与分布对齐紧密耦合,阻碍了优化信号的异步生成、重用和跨模型转移。本文提出代理引导更新信号传输(PUST)框架,将更新信号探索与分布对齐解耦。用轻量级代理模型探索高奖励行为,提取其初始与优化状态间的相对改进信号并传输给主模型引导策略对齐。该解耦流程减少计算开销,支持信号异步生成、缓存和重用,自然支持弱到强的改进及跨模型转移。对Qwen3系列模型在数学和代码领域的系统评估表明,从弱得多的代理中提取的更新信号能稳健且可调地增强更强的主模型。最终,PUST将训练后从单一在线优化过程转变为高度模块化、可重用且高效的范式。

英文摘要

Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward behaviors from policy exploration. While on-policy distillation alleviates this by consolidating independently optimized experts, its reliance on matching absolute expert distributions can yield suboptimal supervision, especially when the target model possesses a different prior or already surpasses the expert's capabilities. To alleviate this, we introduce Proxy OPD (P-OPD), an asynchronous post-training framework that transfers reward-induced policy improvements rather than absolute policy distributions. P-OPD first optimizes a proxy policy via reward feedback. It then extracts the relative distributional changes between the proxy's initial and optimized states, transferring these directional updates through the target model's own on-policy trajectories while retaining the target policy as the reference. This decoupled formulation requires the proxy to provide merely a useful direction of improvement rather than superior absolute capability, enabling update signals from older or weaker proxies to remain highly effective. Systematic experiments on Qwen3-family models across mathematical reasoning and code generation demonstrate that P-OPD consistently enhances already strong target models. Furthermore, transfer intensity can be dynamically modulated through signal scaling, making the extracted update signals seamlessly reusable across diverse model variants and training configurations. These results establish relative policy updates as highly reusable, adjustable assets for scalable, reward-based post-training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑