AI 中文总结
针对在线策略自蒸馏(OPSD)中特权信息依赖参考特定信息的问题,提出 PS-OPSD 方法,在 1.7B-8B 模型规模的三个数学推理基准上,仅基于问题的准确率达最高,验证了其有效性。
AI 中文摘要
在线策略自蒸馏(OPSD)通过利用模型基于参考解的特权视图来监督仅观察问题的学生视图,从而提升推理能力。然而,教师提供的 token 级目标可能依赖于推理时不可用的参考特定信息。我们提出问题空间引导的 OPSD(PS-OPSD),其用描述初始状态、目标条件、约束及选定状态转移路径的基于轨迹的指导,替代完整解。学生 rollout 与 OPSD 目标保持不变。在三个数学推理基准及 17 亿至 80 亿的模型规模范围内,PS-OPSD 在对比方法中取得仅基于问题的最高 aggregate 准确率。受控实验进一步表明,指导相关性与路径一致性促成了这些提升,凸显特权信息的表示是 OPSD 中一项重要的设计选择。
英文摘要
On-policy self-distillation (OPSD) improves reasoning by using a privileged view of a model conditioned on reference solutions to supervise a student view that observes only the question. However, the teacher-provided token-level targets may depend on reference-specific information unavailable at inference time. We propose Problem-Space-Guided OPSD (PS-OPSD), which replaces the complete solution with trajectory-grounded guidance describing the initial state, goal conditions, constraints, and a selected state-transition path. The student rollout and OPSD objective remain unchanged. Across three mathematical reasoning benchmarks and model scales ranging from 1.7B to 8B, PS-OPSD achieves the highest aggregate question-only accuracy among the compared methods. Controlled experiments further indicate that guidance relevance and path coherence contribute to these gains, highlighting the representation of privileged information as an important design choice in OPSD.