arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RESUME:基于运动与残差信号的循环状态更新,用于高效视频语言建模

RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling

Can Zhang, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Dianhai Yu, Ruirui Li

arXiv 2609.39563首次发表:更新:

发表机构

Beijing University of Chemical Technology; Baidu, Inc.(北京化工大学; 百度公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RESUME 提出有状态编解码器表示,以锚定 I 帧初始化潜在状态,预测帧作为更新,通过共享读出生成令牌,提升视频语言模型的时间推理性能。

AI 中文摘要

现有的视频语言模型独立地对采样的 RGB 帧进行编码,因此长视频要么耗尽令牌预算,要么丢失采样帧之间的变化。编解码器感知的前端读取编码产生的运动向量和残差,但在其部署形式中,每个预测帧仍然被单独分词:令牌是当前基元的函数,而不是携带的参考的函数。我们认为更自然的函数是两者的结合——当前基元和携带的参考。一个片段及其时间反转共享相同的帧,仅在变化顺序上有所不同——这是一个对称池化在构造上丢弃的轴,并且在 VideoLM 使用的冻结视觉特征中是非空的——而编解码器循环已经按顺序将这些变化与参考状态组合。我们引入了 RESUME,一种有状态的编解码器表示:一个锚定 I 帧初始化一个紧凑的潜在状态,每个后续的预测帧作为对该状态的更新被消费,一个共享的读出从累积状态中暴露与 VideoLM 兼容的令牌。编解码器预测因此保持在表示层面,并作为轨迹交给语言模型,而不是一组独立的令牌组。在与先前编解码器感知方法相同的每个预测帧令牌预算下,预测帧作为前端已知内容的读出进入语言模型,而不是仅对当前基元的编码。在十个基准测试中,收益集中在时间推理上:在所有三个时间基准上,RESUME 都优于 RGB 帧基线 LLaVA-Video-7B(在 TempCompass、TOMATO 和 MVBench 上分别提高 2.8、5.1 和 3.9 分)和基于编解码器的基线 CoPE-7B,同时在一般和长形式问答上保持竞争力。冻结转换测试进一步显示了锚定依赖性、顺序敏感性以及超出训练范围的有用展开行为。

英文摘要

Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals that encoding produced, but in their deployed form each predictive frame is still tokenized on its own: the tokens are a function of the current primitives, not of a carried reference. We argue that a more natural function is of both---the current primitives and a carried reference. A clip and its time reversal share the same frames and differ only in the order of changes---an axis that symmetric pooling discards by construction, and that is non-empty in the frozen vision features VideoLMs use---and the codec recurrence already composes those changes in order against a reference state. We introduce RESUME, a stateful codec representation: an anchor I-frame initializes a compact latent state, each subsequent predictive frame is consumed as an update to that state, and a shared readout exposes VideoLM-compatible tokens from the accumulated state. Codec prediction is thereby kept at the representation level and handed to the language model as a trajectory, not as a set of independent token groups. At the same per-predictive-frame token budget as prior codec-aware methods, a predictive frame enters the language model as a readout of what the front-end already knows, not as an encoding of the current primitives alone. Across ten benchmarks, the gains concentrate on temporal reasoning: on all three temporal benchmarks RESUME improves over both the RGB-frame baseline LLaVA-Video-7B (by 2.8, 5.1, and 3.9 points on TempCompass, TOMATO, and MVBench) and the codec-based baseline CoPE-7B, while staying competitive on general and long-form QA. Frozen-transition tests further show anchor dependence, order sensitivity, and useful rollout behavior beyond the training horizon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑