arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16316cs.CVcs.AIcs.CL

深度思维对齐:面向视频推理的轨迹级潜在知识蒸馏

Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning

Ao Shen, Yongheng Zhang, Yinghui Li, Manning Wang, Di Yin, Xing Sun

首次发表
浏览论文内容

中文总结 AI 辅助

针对视频推理LMMs的高计算成本问题,提出Latent-OPD方法,通过轨迹级潜在知识蒸馏与渐进式教师前瞻策略提升轻量模型性能,在6个基准上优于仅输出监督的OPD。

中文摘要 AI 辅助

用于视频推理的大型多模态模型(LMMs)长期以来受限于处理海量视觉信息的高计算成本,这一困境促使研究者将大型模型的推理能力迁移至更小、更高效的模型。在线策略蒸馏(OPD)是一种有前景的解决方案,其核心是匹配学生模型生成轨迹上的输出token分布。然而,视频推理通常依赖于多帧累积的证据,在此背景下,输出级监督仅能捕获通过token预测表达的信息,无法直接约束推理过程中形成的潜在表示。为解决这一局限,我们提出Latent-OPD,该方法在OPD基础上引入轨迹级潜在知识蒸馏。具体而言,我们的方法聚焦于每条轨迹末尾的位置,此处的隐藏状态可有效汇总累积的视觉证据与推理上下文。此外,我们引入渐进式教师前瞻策略,使学生模型的中高层层与教师模型的更深层对齐。在6个视频推理基准上的实验表明,Latent-OPD始终优于仅使用输出的OPD;值得注意的是,在帧数有限、长视频或需要复杂证据聚合的场景中,性能提升尤为显著。这些结果证实Latent-OPD是一种高效的帧高效视频推理方法。

英文摘要

Large Multimodal Models (LMMs) for video reasoning have long been hindered by the high computational cost of processing vast amounts of visual information. This dilemma motivates the transfer of the reasoning capabilities of large models to smaller, more efficient ones. On-Policy Distillation (OPD) offers a promising solution by matching output-token distributions along student-generated trajectories. However, video reasoning often depends on evidence accumulated across multiple frames. In this context, output-level supervision only captures information expressed through token predictions and does not directly constrain the latent representations formed during reasoning. To address this limitation, we propose Latent-OPD, which augments OPD with trajectory-level latent distillation. Specifically, our method focuses on the position at the end of each trajectory, where hidden states effectively summarize the accumulated visual evidence and reasoning context. Furthermore, we introduce a progressive teacher-lookahead strategy, which aligns middle-to-late student layers with increasingly deeper teacher layers. Experiments on six video reasoning benchmarks show that Latent-OPD consistently outperforms output-only OPD. Notably, the improvements are particularly pronounced in scenarios with limited frames, long videos, or tasks requiring complex evidence aggregation. These results establish Latent-OPD as a highly effective approach to frame-efficient video reasoning.

↑