发表机构
Stanford University; Vast Intelligence Lab; Google DeepMind(斯坦福大学; Vast Intelligence Lab; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种基于图像表示构建紧凑视频编码器的方法,通过分离逐帧表示、跨帧令牌分配和时间交互,在13个基准上以约28%-35%的视觉令牌匹配全图像性能,并实现2.16倍加速。
AI 中文摘要
视频编码器的设计决定了帧之间何时开始交互,以及哪些特定于帧的视觉证据仍可供语言模型访问。原生视频通路在视觉编码期间将相邻帧耦合,而图像通路则保留独立计算的帧表示,但在所有图像令牌都被转发时会产生更大的视觉令牌成本。我们提出了一个基本问题:是否可以在图像表示的基础上构建紧凑的视频编码器?为了回答这个问题,我们将经常耦合的三个操作分开:逐帧表示、跨帧令牌分配和时间交互。一个冻结的图像编码器首先产生特定于帧的候选。一个查询感知的选择器随后根据相关性、多样性和跨帧对应性在帧间分配固定的令牌预算,之后一个轻量级学习的精炼器读取相邻帧上下文,并仅将残差更新写入保留的锚点。这保留了源位置,并将视觉输出保持在固定预算内。在13个基准测试和三个视觉-语言骨干网络上,所得通路在仅使用约28%-35%的视觉令牌的情况下,匹配了全图像聚合性能。具体来说,在Qwen3-VL-8B上,它使用1,535个视觉令牌实现了13个基准的宏平均62.75,而完整的图像通路在4,424个令牌下为62.58,原生Conv3D在2,212个令牌下为59.49。在Qwen3-VL-32B上,它达到了13个基准的宏平均66.28,而图像通路为66.09,同时提供了2.16倍的端到端加速。这些结果表明,紧凑的视频编码不需要早期的时间混合:可以先保留特定于帧的证据,联合分配,并在选择后进行时间上下文化。
英文摘要
The design of a video encoder determines when frames begin to interact and which frame-specific visual evidence remains accessible to the language model. Native video pathways couple neighboring frames during visual encoding, whereas image pathways preserve independently computed frame representations but incur a much larger visual-token cost when all image tokens are forwarded. We ask a basic question: whether a compact video encoder can instead be built on image representations. To answer this question, we separate three operations that are often coupled: per-frame representation, cross-frame token allocation, and temporal interaction. A frozen image encoder first produces frame-specific candidates. A question-aware selector then allocates a fixed token budget across frames using relevance, diversity, and cross-frame correspondence, after which a lightweight learned refiner reads neighboring-frame context and writes residual updates only to the retained anchors. This preserves source positions and keeps the visual output at the fixed budget. Across 13 benchmarks and three vision-language backbones, the resulting pathway matches full-image aggregate performance while using only about 28%-35% of its visual tokens. Specifically, on Qwen3-VL-8B, it achieves a 13-benchmark macro-average of 62.75 with 1,535 visual tokens, compared with 62.58 for the full Image pathway at 4,424 tokens and 59.49 for native Conv3D at 2,212 tokens. On Qwen3-VL-32B, it reaches a 13-benchmark macro-average of 66.28, compared with 66.09 for Image, while providing a 2.16x end-to-end speedup. These results show that compact video encoding does not require early temporal mixing: frame-specific evidence can be preserved first, allocated jointly, and temporally contextualized after selection.