arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

指导而非回答:在策略内上下文蒸馏中使用指令特权

Instruct, Not Answer: Using Instruction Privileges in On-Policy Context Distillation

Hantao Yu, Xiaoxue Han, Udaya Ghai, Ferhat Erata, Joseph Lilien, Aman Goel, Ali Torkamani

arXiv 2609.32201首次发表:更新:

AI 中文总结

本研究提出在策略内上下文蒸馏中使用通用指令而非实例特定黄金答案作为特权,证明该方法在分布外准确率上大幅提升,同时保持域内性能,且指令简短统一。

AI 中文摘要

策略内上下文蒸馏(OPCD)最近已成为一种强大的技术,用于将上下文转移到学生模型并实现自我改进。在OPCD中,教师以特权信息为条件,目标是最小化特权教师与学生之间在学生生成令牌上评估的Kullback-Leibler(KL)散度。许多现有研究表明,使用实例特定的黄金答案或黄金演示作为默认特权可能会损害训练性能,尤其是在分布外(OOD)情况下。在这项工作中,我们转而设计针对训练样本中观察到的常见学生错误的通用指令,并表明这种简单指令可以胜过黄金作为OPCD特权。在自动形式化任务中,使用匹配的格式指令作为特权可以在OOD准确率上大幅超过黄金。在使用ProverQA、ProofWriter和ProntoQA作为数据集,以及Qwen3-Thinking和Olmo3-Thinking系列作为模型的8个实验中,有7个实验中,匹配指令特权在OOD上比黄金高出4到17个百分点,同时在域内与黄金持平。每条指令只有几句话(因此与所有实例特定的黄金相比包含的信息少得多),并且统一应用于每个训练样本。这些结果表明,一种通用指令,同样适用于源域和目标域示例,在OPCD中可以比实例特定的黄金具有显著更强的可迁移性,同时保持域内性能。

英文摘要

On-Policy Context Distillation (OPCD) has recently emerged as a powerful technique for transferring context to student models and for self-improvement. In OPCD, the teacher is conditioned on privileged information, and the goal is to minimize the Kullback-Leibler (KL) divergence between the privileged teacher and the student, evaluated on student-generated tokens. Many existing studies show that using instance-specific gold answers or gold demonstrations as the default privilege can hurt training performance, especially out-of-distribution (OOD). In this work, we instead design general instructions that target common student mistakes observed on the training samples, and show that such simple instructions can outperform gold as the OPCD privilege. In autoformalization tasks, using a matched formatting instruction as the privilege could outperform gold in OOD accuracy by a large margin. In 7 out of 8 experiments using ProverQA, ProofWriter, and ProntoQA as datasets, and Qwen3-Thinking and Olmo3-Thinking families as models, matched instruction privileges outperform gold in OOD by 4 to 17 points, while remaining on par with gold in-domain. Each instruction is only a few sentences (and thus contains much less information compared to all instance-specific gold) and is applied uniformly to every training sample. These results indicate that a general instruction, which applies equally to source and target domain examples, can be substantially more transferable than instance-specific gold in OPCD while maintaining in-domain performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑