arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-08 至 2025-10-08 共收录 10 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 4 篇

2510.06077 2025-10-08 cs.CV cs.AI 83%

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

Mi Luo, Zihui Xue, Alex Dimakis, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) UC Berkeley(伯克利大学) Bespoke Labs(Bespoke实验室)

专题命中 视频理解 :video reasoning(title,abstract);video understanding(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025, Project page: https://vision.cs.utexas.edu/projects/video-ver/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05836 2025-10-08 cs.CV 83%

Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow

Ruyang Liu, Shangkun Sun, Haoran Tang, Ge Li, Wei Gao

机构 * School of Electronic and Computer Engineering, Shenzhen Graduate School, 2 Peng Cheng LaboratoryPeking University(1 电子与计算机工程学院,深圳研究生院,2 深圳鹏城实验室,北京大学)

专题命中 视频理解 :video understanding(title,abstract);long video(abstract);分类 cs.CV

Comments Accepted to ICCV' 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03501 2025-10-08 cs.CV 83%

LV-MAE: Learning Long Video Representations through Masked-Embedding Autoencoders

Ilan Naiman, Emanuel Ben-Baruch, Oron Anschel, Alon Shoshan, Igor Kviatkovsky, Manoj Aggarwal, Gerard Medioni

机构 * Amazon(亚马逊)

专题命中 视频理解 :long video(title,abstract);video-language(abstract);分类 cs.CV

Comments Accepted to the International Conference on Computer Vision, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15192 2025-10-08 cs.CV 57%

Leveraging Foundation Models for Multimodal Graph-Based Action Recognition

Fatemeh Ziaeetabar, Florentin Wörgötter

机构 * School of Mathematics, Statistics and Computer Science, College of Science, University of Tehran(数学、统计与计算机科学学院,科学学院,塔里斯坦大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 2 篇

2510.06209 2025-10-08 cs.CV 79%

Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models

Jiahao Wang, Zhenpei Yang, Yijing Bai, Yingwei Li, Yuliang Zou, Bo Sun, Abhijit Kundu, Jose Lezama, Luna Yue Huang, Zehao Zhu, Jyh-Jing Hwang, Dragomir Anguelov, Mingxing Tan, Chiyu Max Jiang

机构 * Johns Hopkins University(约翰霍普金斯大学) Waymo Google DeepMind(谷歌DeepMind)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05661 2025-10-08 cs.CV cs.MM 62%

When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach

Daniel Gonzálbez-Biosca, Josep Cabacas-Maso, Carles Ventura, Ismael Benito-Altamirano

机构 * eHealth Center, Faculty of Computer Science, Multimedia and Telecommunications, Universitat Oberta de Catalunya(eHealth中心,计算机科学、多媒体与电信学院,开放大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 2 篇

2501.19252 2025-10-08 cs.CV 85%

Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search

Yuta Oshima, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta

机构 * The University of Tokyo(东京大学) Google DeepMind(谷歌DeepMind)

专题命中 视频扩散模型 :text-to-video(title,abstract);video generation(abstract);video diffusion(abstract);分类 cs.CV

Comments Accepted to NeurIPS2025. Website: https://sites.google.com/view/t2v-dlbs and Code: https://github.com/shim0114/T2V-Diffusion-Search

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21653 2025-10-08 cs.CV 83%

Think Before You Diffuse: Infusing Physical Rules into Video Diffusion

Ke Zhang, Cihan Xiao, Jiacong Xu, Yiqun Mei, Vishal M. Patel

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 长视频与时序推理 2 篇

2510.06040 2025-10-08 cs.CV cs.AI 83%

VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization

Xinye Cao, Hongcan Guo, Jiawen Qian, Guoshun Nan, Chao Wang, Yuqi Pan, Tianhao Hou, Xiaojuan Wang, Yutong Gao

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Minzu University of China(民族大学)

专题命中 长视频与时序推理 :long video(title,abstract);video understanding(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05367 2025-10-08 cs.CV cs.LG 79%

LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation

Yang Xiao, Gen Li, Kaiyuan Deng, Yushu Wu, Zheng Zhan, Yanzhi Wang, Xiaolong Ma, Bo Hui

机构 * University of Tulsa(塔尔萨大学) Clemson University(克莱姆斯大学) The University of Arizona(亚利桑那大学) Northeastern University(东北大学) Microsoft Research(微软研究院)

专题命中 长视频与时序推理 :video generation(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏