arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VideoChat3:用于高效通用视频理解的全开放视频多模态语言模型

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hongjie Zhang, Yi Wang, Yu Qiao, Yali Wang, Ziwei Liu, Kai Chen, Limin Wang

arXiv 2607.14935首次发表:更新:

发表机构

Nanjing University; Shanghai AI Laboratory; Nanyang Technological University; Peking University(南京大学; 上海人工智能实验室; 南洋理工大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视频理解开源模型局限,提出VideoChat3。通过I3D-ViT等提升效率,利用可扩展视频数据合成管道生成训练数据集提升泛化性,以4B参数在多基准测试中超越同等或更多参数的开源模型,实现泛化与计算效率平衡。

AI 中文摘要

视频理解领域虽有进展,但当前开源模型存在局限。它们难以跨多种视频类型泛化,计算需求高限制了效率与可扩展性,且大多模型部分开放。为此引入全开放、高效且通用的以视频为中心的多模态语言模型VideoChat3。通过Inflated 3D Vision Transformer (I3D-ViT)等设计提升效率,开发可扩展视频数据合成管道生成三个训练数据集提升泛化性,实验表明其在泛化和计算效率上取得平衡,超越了同等或更多参数的开源模型。

英文摘要

Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-driven development. To address these issues, we introduce VideoChat3, a fully open, efficient, and generalist video-centric MLLM. VideoChat3 advances video understanding through two complementary designs. For efficiency, we introduce Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for Streaming Video Perception, which enables efficient spatiotemporal representation and reduces the cost of processing video inputs during training and inference. For effectiveness, we develop a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K, covering general, long-form, and streaming video scenarios, improving the model's generalization across domains. By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑