arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在褪色之前:在推理时强化VideoLLMs中的时间表示

Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs

Youngwoo Shin, Yusung Ro, Minseo Kim, Junmo Kim

arXiv 2610.01595首次发表:更新:

发表机构

Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现VideoLLMs时间推理弱源于中间层时间信息向输出层逐渐褪色,提出无需训练的时间激活注入(TAI)方法,在峰值处提取并重新注入表示,显著提升三个模型四个基准的时间推理性能。

AI 中文摘要

视频大语言模型(VideoLLMs)按顺序接收帧,并解释视觉内容如何沿时间轴演变,然而时间推理在不同架构中仍然是一个持续存在的弱点。反转视频的帧顺序,这种变换本应使时间答案反转,却常常使最终预测保持不变。我们通过定义时间发散向量τ_l(即由反转时间顺序引起的逐层表示差异)来调查这一失败源于何处。追踪其幅度在各层间的变化,揭示了一个一致的时间发散轮廓,其中发散在中间层达到峰值,并逐渐向输出层减弱。我们确认这一峰值是时间推理特有的,并且对预测具有功能关键性,从而确定VideoLLMs在中间层获取时间信息,但未能将其保持到输出层。这种逐渐褪色的现象启发了我们的方法——时间激活注入(TAI),该方法为每个输入在轮廓峰值处提取τ_l,并根据测得的衰减将其重新注入后续层。TAI无需训练,并在三个VideoLLMs和四个基准上持续改善时间推理,对非时间任务的影响可忽略不计。代码可在以下网址获取:此https URL。

英文摘要

Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector $τ_l$, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts $τ_l$ at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at https://github.com/Youngwoo-git/Before-It-Fades.

CommentsAccepted to NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑