AI 中文总结
【一句话总结】研究针对策略蒸馏在流模型中的扩展探索不足的问题,提出FlowCTS,通过匹配轨迹等方法得出速度匹配上限并离散化,在多参考设置下有优势,还揭示了基于KL的策略蒸馏的不足,且在策略外设置也表现出色。
AI 中文摘要
虽然策略蒸馏有效地解决了大语言模型训练后稀疏奖励和曝光偏差问题,但其在流模型中的扩展仍未得到充分探索。为此,我们提出了流连续轨迹监督(FlowCTS),它匹配从相同学生访问状态初始化的后续学生轨迹和参考轨迹。利用轨迹和速度场之间的积分关系,我们推导出一个时间加权速度匹配上限,并将其离散化为由监督步数参数化的实际目标。在多参考设置下,单状态FlowCTS-OPD比基于香草KL的OPD收敛更快。FlowCTS-OPD将GenEval从0.90提高到0.93,将OCR从0.90提高到0.92,将PickScore从22.75提高到23.06,同时在所有目标指标上优于混合奖励RL基线。进一步分析揭示了基于香草KL的OPD中由于其辅助SDE过渡核而产生的明显时间监督不匹配。除了策略设置外,FlowCTS也始终优于香草SFT,特别是在OCR上,而增加监督步数在更丰富的轨迹信息和更大的优化难度之间表现出权衡。
英文摘要
While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored. To this end, we propose Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and reference trajectories initialized from the same student-visited state. Using the integral relation between trajectories and velocity fields, we derive a temporally weighted velocity-matching upper bound and discretize it into practical objectives parameterized by the number of supervision steps. Under a multi-reference setup, single-state FlowCTS-OPD outperforms vanilla KL-based OPD with faster convergence. FlowCTS-OPD improves GenEval from 0.90 to 0.93, OCR from 0.90 to 0.92, and PickScore from 22.75 to 23.06, while outperforming a mixed-reward RL baseline across all target metrics. Further analysis reveals a clear temporal supervision mismatch in vanilla KL-based OPD arising from its auxiliary SDE transition kernels. Beyond on-policy setting,FlowCTS also consistently outperforms vanilla SFT , particularly on OCR, while increasing supervision steps exhibit a trade-off between richer trajectory information and greater optimization difficulty.