SAPD:步骤对齐的特权蒸馏
SAPD: Step-Aligned Privileged Distillation
浏览论文内容
中文总结 AI 辅助
提出步骤对齐特权蒸馏(SAPD),一种无需轨迹生成的自蒸馏方法,通过将演示数据转化为步骤对齐的分布监督,在数学推理上优于监督微调,且与在策略方法竞争力相当,并实现约2倍训练加速。
中文摘要 AI 辅助
在策略上的后训练可以通过从模型自身的轨迹中学习来改进大型语言模型,但需要昂贵的轨迹生成成本。我们探究固定的演示数据是否能够通过更好的监督来支持具有竞争力的离策略学习。我们的前提是,这些演示数据的效用不仅取决于训练轨迹本身,还取决于监督是否在延续候选中提供了有信息量的偏好,并将这种指导与正在学习的推理决策联系起来。我们引入了步骤对齐特权蒸馏(SAPD),这是一种无需轨迹生成的自蒸馏方法,它将演示数据转化为步骤对齐的分布监督。其关键洞察在于利用参考解决方案的已知进展,将每个推理转换与有针对性的特权指导相关联,而不是将解决方案视为无差别的上下文。在数学推理基准上,SAPD在平均性能上优于监督微调和标签平滑,同时与在策略强化学习和自蒸馏方法保持竞争力。分析支持了上下文相关的分布指导的价值以及将特权信息与当前步骤对齐的益处。SAPD还大致保留了领域外的编码性能,并相对于在策略基线实现了约2倍的训练循环加速。这些发现表明,精心构建的监督可以使完全离策略的后训练成为一种具有竞争力和计算效率的替代方案。我们的代码可在以下网址获取:https://this https URL。
英文摘要
On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation. We ask whether fixed demonstrations can support competitive off-policy learning through better supervision. Our premise is that their usefulness depends not only on the training trajectories, but also on whether supervision provides informative preferences among continuations and connects this guidance to the reasoning decision being learned. We introduce Step-Aligned Privileged Distillation (SAPD), a rollout-free self-distillation method that turns demonstrations into step-aligned distributional supervision. Its key insight is to use the known progression of a reference solution to associate each reasoning transition with targeted privileged guidance, rather than treating the solution as undifferentiated context. On mathematical reasoning benchmarks, SAPD outperforms supervised fine-tuning and label smoothing on average while remaining competitive with on-policy reinforcement learning and self-distillation. Analyses support both the value of context-dependent distributional guidance and the benefit of aligning privileged information with the current step. SAPD also largely preserves out-of-domain coding performance and achieves approximately 2x training-loop speedups over the on-policy baselines. These findings suggest that carefully constructed supervision can make fully off-policy post-training a competitive and computationally efficient alternative. Our code is available at https://github.com/Miaow-Lab/SAPD.
发表机构
- City University of Hong Kong(香港城市大学)
- Hong Kong Institute of AI for Science, City University of Hong Kong(香港城市大学香港人工智能科学研究院)
- Shenzhen University(深圳大学)
机构由 AI 辅助整理,请以论文原文为准。