发表机构
Peking University; Kling Team; Tsinghua University; Shanghai Jiao Tong University; Zhongguancun Academy(北京大学; KLING团队; 清华大学; 上海交通大学; 中关村学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对开放域大语言模型训练缺乏可验证奖励的问题,提出Flux-OPD范式,利用演化上下文作为监督,通过分解反向KL目标优化,在开放域任务上性能优于现有OPD范式。
AI 中文摘要
在开放域中训练大语言模型时,缺乏可验证的奖励,难以将任务偏好形式化为有效监督。上下文可传递此类偏好,但将其蒸馏到学生模型后便几乎无法提供额外监督,因此需要随学生模型性能演化的上下文。然而,直接将演化上下文用作训练中监督会导致蒸馏目标不稳定以及分布冲突,需要机制来稳定目标并降低冲突权重。本文通过分解反向KL目标分析上下文的影响,得出两项发现:学生模型被蒸馏至上下文条件教师的几何均值,且该目标包含衡量这些教师间冲突的冲突项。基于此分解,本文提出Flux-OPD,这一在线策略蒸馏(OPD)范式将演化上下文用作训练中监督,以捕捉开放域中的任务偏好。Flux-OPD将上下文条件教师与无上下文教师之间的差异视为上下文差异信号,将其作为上下文修正注入无上下文教师锚点,并以冲突项为指标对修正强度加权。在开放域任务上的实验表明,Flux-OPD优于现有OPD范式,凸显了将教师监督与演化上下文结合的潜力。
英文摘要
Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize target and downweight conflicts. In this paper, we analyze the effect of contexts through a decomposition of the reverse KL objective, revealing two findings: the student is distilled toward the geometric mean of context-conditioned teachers, and the objective contains a conflict term that measures conflicts among these teachers. Based on this decomposition, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains. Flux-OPD treats the differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher anchor, and weights their correction strength using the conflict term as an indicator. Experiments on open-ended tasks show that Flux-OPD outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.