arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39924cs.CVcs.AI

CoVisco:面向统一图像-视频理解的原生码率视觉编码器与原生令牌压缩

CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

Yulong Liu, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Guibo Zhu, Sirui Han, Dianhai Yu

首次发表
浏览论文内容

中文总结 AI 辅助

CoVisco通过码率原生编码与分段注意力实现原生令牌压缩,在64帧四段设置下仅用400个视觉令牌达到接近或超越OneVision-Encoder的视频理解性能,同时保留细粒度证据。

中文摘要 AI 辅助

视觉-语言模型面临一个基本的扩展瓶颈:视觉令牌的数量随时间和空间分辨率增长,这使得长视频理解对视觉编码器和语言模型而言成本高昂。现有方法通常在密集编码后压缩视觉令牌,导致训练期间使用的表示与部署时所需的紧凑接口不匹配。我们提出CoVisco,一种具有原生令牌压缩的码率原生视觉编码器,用于统一的图像-视频理解。通过结合码率原生输入支持与分段注意力,CoVisco能够在单次前向传播中编码长视觉输入,而无需在所有帧间形成密集的补丁到补丁交互。每个时间分段配备可学习的抽象令牌,学习紧凑的分段级表示,而细粒度的补丁令牌在整个编码器中保持可用。交替的段内和抽象通信层通过抽象令牌通道保留视频级上下文。一个轻量级选择器进一步暴露仅抽象令牌或抽象令牌加上运行时选择的补丁令牌子集,产生紧凑的视觉接口,减少下游多模态大语言模型的视觉上下文和预填充负担,同时在需要时保留细粒度证据。在5.65亿图像-文本对和640万视频上使用对比目标进行预训练,CoVisco在面向视频的嵌入和多模态理解基准上表现出竞争力。在评估的四段、64帧设置中,仅抽象推理仅使用400个视觉令牌,同时实现接近且在某些基准上超过OneVision-Encoder的视频理解性能。选定的补丁令牌进一步改善细粒度视频推理。项目网址:此HTTPS URL。

英文摘要

Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a codec-native vision encoder with native token compression for unified image-video understanding. By combining codec-native input support with segmented attention, CoVisco can encode long visual inputs in a single forward pass without forming dense patch-to-patch interactions across all frames. Each temporal segment is equipped with learnable abstract tokens that learn a compact segment-level representation, while fine-grained patch tokens remain available throughout the encoder. Alternating intra-segment and abstract-communication layers preserve video-level context through the abstract-token channel. A lightweight selector further exposes either abstract tokens alone or abstract tokens augmented with a runtime-selected subset of patch tokens, yielding a compact visual interface that reduces the visual context and prefill burden of downstream MLLMs while retaining fine-grained evidence when needed. Pretrained with contrastive objectives on 565M image--text pairs and 6.4M videos, CoVisco shows competitive performance on video-oriented embedding and multimodal understanding benchmarks. In the evaluated four-segment, 64-frame setting, abstract-only inference uses only 400 visual tokens while achieving video-understanding performance close to, and on some benchmarks exceeding, OneVision-Encoder. Selected patch tokens further improve fine-grained video reasoning. Project URL: https://github.com/ernie-research/CoVisco.git

发表机构

  • ERNIE Team, Baidu Inc.(百度公司ERNIE团队)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑