arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-15 至 2025-10-15 共收录 10 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 2 篇

2510.12160 2025-10-15 cs.CV 79%

State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding

Jiahuan Zhou, Kai Zhu, Zhenyu Cui, Zichen Liu, Xu Zou, Gang Hua

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) the Huazhong University of Science and Technology(华中科技大学) Amazon.com, Inc(亚马逊公司)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11907 2025-10-15 cs.CV 57%

Task-Specific Dual-Model Framework for Comprehensive Traffic Safety Video Description and Analysis

Blessing Agyei Kyem, Neema Jakisa Owor, Andrews Danyo, Joshua Kofi Asamoah, Eugene Denteh, Tanner Muturi, Anthony Dontoh, Yaw Adu-Gyamfi, Armstrong Aboah

机构 * North Dakota State University(北达科塔州立大学) University of Missouri–Columbia(密苏里大学哥伦比亚分校) University of Memphis(孟菲斯大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments This paper was accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 3 篇

2508.03334 2025-10-15 cs.CV 88%

Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation

Xunzhi Xiang, Yabo Chen, Guiyu Zhang, Zhongyu Wang, Zhe Gao, Quanming Xiang, Gonghu Shang, Junqi Liu, Haibin Huang, Yang Gao, Chi Zhang, Qi Fan, Xuelong Li

机构 * Nanjing University(南京大学) Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院) Shanghai Jiao Tong University(上海交通大学) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 视频生成 :video generation(title,abstract);long video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12283 2025-10-15 cs.CV 57%

Dual Learning with Dynamic Knowledge Distillation and Soft Alignment for Partially Relevant Video Retrieval

Jianfeng Dong, Lei Huang, Daizong Liu, Xianke Chen, Xun Yang, Changting Lin, Xun Wang, Meng Wang

机构 * College of Computer and Information Engineering, Zhejiang Gongshang University(浙江工商大学计算机与信息工程学院) Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院) Binjiang Institute of Zhejiang University(浙江大学滨江学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院)

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12231 2025-10-15 cs.CV 57%

BIGFix: Bidirectional Image Generation with Token Fixing

Victor Besnier, David Hurych, Andrei Bursuc, Eduardo Valle

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 1 篇

2510.12785 2025-10-15 cs.CV cs.AI cs.GR 83%

MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars

Felix Taubner, Ruihang Zhang, Mathieu Tuli, Sherwin Bahmani, David B. Lindell

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) Vector Institute Canada(加拿大向量研究所) University of Toronto Canada(多伦多大学加拿大)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments 18 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频问答 1 篇

2510.12299 2025-10-15 cs.IR 50%

An Empirical Study for Representations of Videos in Video Question Answering via MLLMs

Zhi Li, Yanan Wang, Hao Niu, Julio Vizcarra, Masato Taya

专题命中 视频问答 :long video(abstract)

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 长视频与时序推理 2 篇

2510.12422 2025-10-15 cs.CV 88%

VideoLucy: Deep Memory Backtracking for Long Video Understanding

Jialong Zuo, Yongtai Deng, Lingdong Kong, Jingkang Yang, Rui Jin, Yiwei Zhang, Nong Sang, Liang Pan, Ziwei Liu, Changxin Gao

机构 * National Key Laboratory of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(多谱段信息智能处理国家级重点实验室,人工智能与自动化学院,华中科技大学) NUS(国立大学新加坡) S-Lab, NTU(NTU的S实验室) Shanghai AI Lab(上海人工智能实验室)

专题命中 长视频与时序推理 :video understanding(title,abstract);long video(title,abstract);分类 cs.CV

Comments NeurIPS-2025 Accepted Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22622 2025-10-15 cs.CV 88%

LongLive: Real-time Interactive Long Video Generation

Shuai Yang, Wei Huang, Ruihang Chu, Yicheng Xiao, Yuyang Zhao, Xianbang Wang, Muyang Li, Enze Xie, Yingcong Chen, Yao Lu, Song Han, Yukang Chen

机构 * NVIDIA MIT(麻省理工学院) HKUST(GZ)(香港科技大学(广州)) HKU(香港大学) THU(清华大学)

专题命中 长视频与时序推理 :video generation(title,abstract);long video(title,abstract);分类 cs.CV

Comments Code, model, and demos are available at https://github.com/NVlabs/LongLive

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 其他视频模型 1 篇

2505.12434 2025-10-15 cs.CV 79%

VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning

Qi Wang, Yanrui Yu, Ye Yuan, Rui Mao, Tianfei Zhou

机构 * Beijing Institute of Technology(北京理工大学) Shenzhen University(深圳大学)

专题命中 其他视频模型 :video reasoning(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025. Code: https://github.com/QiWang98/VideoRFT

详情

展开后加载摘要…

URL PDF HTML 收藏