AI 中文总结
针对扩散视觉语言模型后训练中掩码与模型揭示决策不一致的问题,提出CT-OPD,结合教师完整响应与学生轨迹掩码重建状态进行在策略蒸馏,在多个基准上平均提升最高9.80分。
AI 中文摘要
扩散视觉语言模型通过逐步解析被掩码的标记来生成答案,这使得在部分解析状态下进行准确的条件预测成为后训练的核心。对已完成的答案进行掩码可以产生连贯的上下文和目标,但预设的掩码并不能反映模型的揭示决策。模型的轨迹捕捉了这些决策,然而其临时的可见标记可能与目标响应相冲突。基于结果的强化学习遵循这些轨迹,但仅提供响应级别的反馈,当采样奖励相同时会失去对比度。为了将连贯的标记级监督与模型的揭示决策对齐,我们提出了反事实轨迹在策略蒸馏(CT-OPD),该方法将完整的教师响应与当前学生的轨迹掩码相结合。CT-OPD在学生的词汇表中对每个教师响应进行重新分词,并在学生反向过程的各个阶段提取未解析位置的掩码。对于每个掩码,它丢弃临时的滚动值,并从教师端点重建部分状态,从而使被监督的位置遵循当前轨迹,同时可见上下文和目标与同一响应保持一致。学生使用其原生的分类损失在这些重建状态上进行训练,并且随着模型的演化,轨迹会不断刷新。在密集和稀疏扩散架构中,CT-OPD持续增强多模态理解和推理能力,在九个基准的平均值上最高提升9.80分。在统一的理解与生成架构上,它还同时改善了视觉理解和图像生成,表明相同的原理可以跨架构和模态迁移。消融实验进一步将这些增益归因于连贯的重建和当前模型的轨迹掩码。
英文摘要
Diffusion vision-language models generate answers by gradually resolving masked tokens, making accurate conditional prediction in partially resolved states central to post-training. Masking completed answers yields coherent contexts and targets, but prescribed masks do not reflect the model's reveal decisions. Its trajectories capture these decisions, yet their provisional visible tokens can conflict with the target response. Outcome-based reinforcement learning follows these trajectories but provides only response-level feedback, which loses contrast when sampled rewards tie. To align coherent token-level supervision with the model's reveal decisions, we introduce Counterfactual Trace On-Policy Distillation (CT-OPD), which combines completed teacher responses with trajectory masks from the current student. CT-OPD retokenizes each teacher response in the student's vocabulary and extracts unresolved-position masks at successive stages of the student's reverse process. For each mask, it discards provisional rollout values and reconstructs the partial state from the teacher endpoint, so the supervised positions follow the current trajectory while the visible context and targets remain consistent with the same response. The student is trained on these reconstructed states with its native categorical loss, and trajectories are refreshed as the model evolves. Across dense and sparse diffusion architectures, CT-OPD consistently enhances multimodal understanding and reasoning capabilities, with gains of up to 9.80 points on the nine-benchmark average. On the unified understanding-and-generation architecture, it also improves both visual understanding and image generation, showing that the same principle transfers across architectures and modalities. Ablations further attribute these gains to coherent reconstruction and current-model trajectory masks.