arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReDraft,而非仅蒸馏:面向持续VLLM后训练的参考驱动修订

ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training

Zhihao Zhang, Mingqi Wu, Qiaole Dong, Enyu Zhou, Shuo Li, Boyang Liu, Jiazheng Zhang, Honglin Guo, Xin Guo, Shaofan Liu, Junzhe Wang, Dingwei Zhu, Minlong Peng, Yuan Hua, Zhiheng Xi, Qi Zhang, Tao Gui, Xuanjing Huang

arXiv 2609.16639首次发表:更新:

发表机构

Fudan University; Shanghai Artificial Intelligence Laboratory; Tencent(复旦大学; 上海人工智能实验室; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ReDraft方法,通过让模型基于专家参考修订自身错误输出并验证后微调,在持续多模态后训练中同时提升新任务性能并显著减少遗忘。

AI 中文摘要

大型多模态模型的持续后训练应在保留预训练能力的同时增加新能力,而这两个目标方向相反。SFT提供显式的目标监督,能从接近零的准确率学习任务,但其离策略目标使模型偏离过远,导致遗忘;RLVR和自蒸馏等在线策略方法保持策略接近性,但当策略尚无法解决任务时提供的信号很少。我们提出ReDraft(参考驱动修订与微调),该方法从模型自身的失败中同时获得两者:仅将专家响应作为参考,让模型修订自身不正确的输出,仅当验证器接受修订时才保留,并对保留的样本进行微调。因此,每个保留的目标都是显式的,但仍接近当前策略。在Qwen2.5-VL-3B/7B上的计数、时钟读取和拼图任务中(其中两个任务准确率接近零),ReDraft在目标任务上比SFT的52.9分高出56.9分,同时将先前任务损失从16.6分降至1.5分(遗忘减少11.3倍),并在两个维度上优于OPSD(提升19.3分,损失6.2分)。数据和参数空间分析符合设计:修订后的目标在基础模型下更可能,且它们引起的更新保持紧凑,比OPSD更紧密地遵循SFT的方向。修复模型自身的输出,而非用专家的输出替换,是让一个目标同时实现两者的关键。

英文摘要

Continual post-training of large multimodal models should add new capabilities while preserving those from pre-training, and the two goals pull in opposite directions. SFT gives explicit target supervision that learns a task from near-zero accuracy, but its off-policy targets move the model far enough to cause forgetting; on-policy methods such as RLVR and self-distillation preserve policy proximity yet supply little signal when the policy cannot yet solve the task. We introduce ReDraft (Reference-Driven Revision and Fine-Tuning), which obtains both from the model's own failures: using an expert response only as a reference, it has the model revise its own incorrect rollout, keeps the revision only if a verifier accepts it, and fine-tunes on what survives. Each retained target is therefore explicit, yet still close to the current policy. Across Counting, Clock Reading, and Jigsaw on Qwen2.5-VL-3B/7B, two of them with near-zero accuracy, ReDraft gains 56.9 points on the target task against SFT's 52.9 while cutting prior-task loss from 16.6 to 1.5 points (11.3x less forgetting), and improves on OPSD along both axes (19.3 gain, 6.2 loss). Data- and parameter-space analyses match the design: revised targets are more probable under the base model, and the updates they induce stay compact and follow SFT's direction more closely than OPSD's. Together, these results show that revising the model's own rollout rather than directly imitating an expert trajectory can reconcile cold-start acquisition with prior-capability retention.

Comments40 pages, 17 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑