发表机构
UCAS; Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd; BJTU; BUAA; ZJU(中国科学院大学; 星辰AGI实验室,中国电信人工智能科技(北京)有限公司; 北京交通大学; 北京航空航天大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MeSD提出多证据自蒸馏框架,通过三个共享参数的证据条件教师和验证引导优化,为VideoLLM提供密集token级监督,在多个视频基准上超越强化学习和自蒸馏基线。
AI 中文摘要
虽然带有可验证奖励的强化学习为VideoLLM提供了可靠的结果监督,但序列级奖励提供的token级指导有限。在策略自蒸馏通过让自教师基于特权信息提供密集的token级监督来解决这一限制。然而,在单一教师上下文中聚合异构证据会模糊跨证据的一致性和冲突。另一个挑战在于确定教师指导应改进基于奖励的更新还是为失败的轨迹提供纠正性监督。为解决这些问题,我们提出了MeSD,一个用于VideoLLM的多证据自蒸馏框架。MeSD构建了三个共享参数、以证据为条件的教师,将真实答案作为共同语义上下文,同时分别纳入时间和空间证据。给定相同的学生生成前缀,MeSD评估相对于答案教师的证据特定偏好,并通过门控的教师特定残差融合教师共同偏好。此外,MeSD引入了验证引导优化,将轨迹分类为成功、失败或不确定。对于成功和不确定的轨迹,MeSD细化token级优势幅度,同时保留奖励派生的符号。对于包含所需证据的已验证失败轨迹,MeSD应用失败条件蒸馏,使用反向KL校正朝融合分布方向。在多个视频基准上的实验表明,与强化学习和自蒸馏基线相比,MeSD持续取得改进。
英文摘要
While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-distillation addresses this limitation by conditioning a self-teacher on privileged information to provide dense token-level supervision. However, aggregating heterogeneous evidence within a single teacher context obscures cross-evidence agreement and conflict. A further challenge lies in determining whether teacher guidance should refine reward-based updates or provide corrective supervision for failed trajectories. To address these issues, we propose MeSD, a multi-evidence self-distillation framework for VideoLLMs. MeSD constructs three evidence-conditioned teachers with shared parameters, using the ground-truth answer as a common semantic context while separately incorporating temporal and spatial evidence. Given the same student-generated prefixes, MeSD evaluates evidence-specific preferences relative to the Answer Teacher and fuses teacher-common preferences with gated teacher-specific residuals. Furthermore, MeSD introduces Verification-Guided Optimization to classify trajectories as Success, Failure, or Indeterminate. For Success and Indeterminate trajectories, MeSD refines token-level advantage magnitudes while preserving reward-derived signs. For verified failure trajectories that contain the required evidence, MeSD applies failure-conditioned distillation, using reverse-KL correction toward the fused distribution. Experiments on multiple video benchmarks demonstrate consistent gains over reinforcement learning and self-distillation baselines.