发表机构
Beijing University of Chemical Technology; Baidu, Inc.(北京化工大学; 百度公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RESUME 提出有状态编解码器表示,以锚定 I 帧初始化潜在状态,预测帧作为更新,通过共享读出生成令牌,提升视频语言模型的时间推理性能。
AI 中文摘要
现有的视频语言模型独立地对采样的 RGB 帧进行编码,因此长视频要么耗尽令牌预算,要么丢失采样帧之间的变化。编解码器感知的前端读取编码产生的运动向量和残差,但在其部署形式中,每个预测帧仍然被单独分词:令牌是当前基元的函数,而不是携带的参考的函数。我们认为更自然的函数是两者的结合——当前基元和携带的参考。一个片段及其时间反转共享相同的帧,仅在变化顺序上有所不同——这是一个对称池化在构造上丢弃的轴,并且在 VideoLM 使用的冻结视觉特征中是非空的——而编解码器循环已经按顺序将这些变化与参考状态组合。我们引入了 RESUME,一种有状态的编解码器表示:一个锚定 I 帧初始化一个紧凑的潜在状态,每个后续的预测帧作为对该状态的更新被消费,一个共享的读出从累积状态中暴露与 VideoLM 兼容的令牌。编解码器预测因此保持在表示层面,并作为轨迹交给语言模型,而不是一组独立的令牌组。在与先前编解码器感知方法相同的每个预测帧令牌预算下,预测帧作为前端已知内容的读出进入语言模型,而不是仅对当前基元的编码。在十个基准测试中,收益集中在时间推理上:在所有三个时间基准上,RESUME 都优于 RGB 帧基线 LLaVA-Video-7B(在 TempCompass、TOMATO 和 MVBench 上分别提高 2.8、5.1 和 3.9 分)和基于编解码器的基线 CoPE-7B,同时在一般和长形式问答上保持竞争力。冻结转换测试进一步显示了锚定依赖性、顺序敏感性以及超出训练范围的有用展开行为。
英文摘要
Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals that encoding produced, but in their deployed form each predictive frame is still tokenized on its own: the tokens are a function of the current primitives, not of a carried reference. We argue that a more natural function is of both---the current primitives and a carried reference. A clip and its time reversal share the same frames and differ only in the order of changes---an axis that symmetric pooling discards by construction, and that is non-empty in the frozen vision features VideoLMs use---and the codec recurrence already composes those changes in order against a reference state. We introduce RESUME, a stateful codec representation: an anchor I-frame initializes a compact latent state, each subsequent predictive frame is consumed as an update to that state, and a shared readout exposes VideoLM-compatible tokens from the accumulated state. Codec prediction is thereby kept at the representation level and handed to the language model as a trajectory, not as a set of independent token groups. At the same per-predictive-frame token budget as prior codec-aware methods, a predictive frame enters the language model as a readout of what the front-end already knows, not as an encoding of the current primitives alone. Across ten benchmarks, the gains concentrate on temporal reasoning: on all three temporal benchmarks RESUME improves over both the RGB-frame baseline LLaVA-Video-7B (by 2.8, 5.1, and 3.9 points on TempCompass, TOMATO, and MVBench) and the codec-based baseline CoPE-7B, while staying competitive on general and long-form QA. Frozen-transition tests further show anchor dependence, order sensitivity, and useful rollout behavior beyond the training horizon.