交互式策略蒸馏与双向提出-验证机制
Interactive-Policy Distillation with Bidirectional Propose-and-Verify
浏览论文内容
中文总结 AI 辅助
提出交互式策略蒸馏(IPD),通过双向提出-验证机制和来源分离损失,解决在线策略蒸馏中教师失锚问题,提升数学推理任务性能与数据效率。
中文摘要 AI 辅助
在线策略蒸馏(OPD)在学生模型自生成的轨迹上,利用密集的令牌级教师反馈来训练学生模型。然而,朴素的OPD可能遭受教师失锚问题,即学生的推理轨迹偏离教师过远,导致教师被查询到其几乎不会访问的状态,从而提供不可靠的监督。我们提出交互式策略蒸馏(IPD),该方法对学生 rollout 应用自适应教师干预。在双向提出-验证状态机下,学生和教师交替扮演提出者和验证者的角色,协作生成混合来源轨迹。然后根据每个令牌的来源应用不同的监督。这种双向提出-验证机制和来源分离损失使IPD不仅成为一种更高效的蒸馏方法,也成为在线和离线策略范式之间的统一桥梁。为了使交错的双模型 rollout 更高效,我们还设计了一个专用的融合推理引擎,该引擎在一个服务实例中同时托管两个模型,并配有独立的KV缓存,同时实例化状态机模型来分发、收集和处理请求。在数学推理任务和多个师生模型对上,使用IPD训练的学生模型不仅优于使用OPD训练的模型,还表现出更高的数据效率。具体而言,当将Qwen3-30B-A3B蒸馏到Qwen3-1.7B-Base时,与OPD相比,IPD在基准平均准确率上带来了+3.28的mean@8和+3.28的best@8提升。此外,IPD仅消耗约1/4的训练样本和步骤,就能超越在整个训练数据集上训练一个epoch的OPD。我们还研究了不同损失变体和接管/交回配置的影响,并展示了IPD在不同训练数据上的鲁棒性。
英文摘要
On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD may suffer from teacher unanchoring, where the student's reasoning trajectory drifts far from the teacher, causing the teacher to be queried on states it would hardly visit and thus provide unreliable supervision. We propose Interactive-Policy Distillation (IPD), which applies adaptive teacher intervention to the student rollout. Under a bidirectional propose-and-verify state machine, the student and teacher alternately exchange their roles as proposer and verifier, and collaboratively generate mixed-source trajectories. Then different supervisions are applied according to the source of each token. This bidirectional propose-and-verify mechanism and the source-split loss make IPD not only a more performant distillation method, but also a unified bridge between on-policy and off-policy paradigms. To make the interleaved dual-model rollouts more efficient, we also design a dedicated fused inference engine that co-hosts both models in one serving instance with separate KV caches and instantiates the state machine model to distribute, collect, and process requests. On math reasoning tasks and across multiple teacher-student model pairs, student models trained with IPD not only outperform those trained with OPD, but also demonstrate higher data efficiency. Specifically, when distilling Qwen3-30B-A3B into Qwen3-1.7B-Base, IPD brings a +3.28 mean@8 and a +3.28 best@8 benchmark-averaged accuracy improvement compared with OPD. Besides, IPD only consumes about 1/4 of the training examples and steps to outperform OPD trained on the whole training dataset for one epoch. We also investigate the impact of different loss variants and takeover / handback configurations, and demonstrate the robustness of IPD on different training data.
发表机构
- University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
- Morgan Stanley(摩根士丹利)
- Rutgers University–New Brunswick(罗格斯大学新布朗斯维克分校)
机构由 AI 辅助整理,请以论文原文为准。