arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-31 至 2025-10-31 共收录 5 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1 篇

2510.26027 2025-10-31 cs.CV 57%

Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders

Ali Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk, Mohsen Fayyaz

机构 * Leibniz University Hannover(莱比锡大学汉诺威分校) L3S Research Center(L3S研究中心) Microsoft(微软公司)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 1 篇

2510.24134 2025-10-31 cs.CV cs.AI cs.CL 88%

VC4VG: Optimizing Video Captions for Text-to-Video Generation

Yang Du, Zhuoran Lin, Kaiqiang Song, Biao Wang, Zhicheng Zheng, Tiezheng Ge, Bo Zheng, Qin Jin

机构 * School of Information, Renmin University of China(中国人民大学信息学院) Taobao & Tmall Group of Alibaba(阿里巴巴淘宝与天猫集团)

专题命中 视频生成 :video generation(title,abstract);text-to-video(title,abstract);分类 cs.CV

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 2 篇

2411.10501 2025-10-31 cs.CV cs.LG 83%

OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models

Mathis Koroglu, Hugo Caselles-Dupré, Guillaume Jeanneret Sanmiguel, Matthieu Cord

机构 * Obvious Research ISIR - Sorbonne University(ISIR - 索邦大学)

专题命中 视频扩散模型 :video diffusion(title);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments 8 pages, 1 supplementary page, 9 figures

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2025, pp. 6225-6235

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08325 2025-10-31 cs.CV 70%

GameFactory: Creating New Games with Generative Interactive Videos

Jiwen Yu, Yiran Qin, Xintao Wang, Pengfei Wan, Di Zhang, Xihui Liu

机构 * The University of Hong Kong(香港大学) Kuaishou Technology(快手科技)

专题命中 视频扩散模型 :video generation(abstract);video diffusion(abstract);分类 cs.CV

Comments ICCV 2025 Highlight, Project Page: https://yujiwen.github.io/gamefactory

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频数据与评测 1 篇

2510.26802 2025-10-31 cs.CV cs.AI cs.CL 57%

Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark

Ziyu Guo, Xinyan Chen, Renrui Zhang, Ruichuan An, Yu Qi, Dongzhi Jiang, Xiangtai Li, Manyuan Zhang, Hongsheng Li, Pheng-Ann Heng

机构 * CUHK(香港中文大学) IMIXR MMLab Peking University(北京大学) Northeastern University(东北大学)

专题命中 视频数据与评测 :video generation(abstract);分类 cs.CV

Comments Project Page: https://video-cof.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏