arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-17 至 2025-10-17 共收录 13 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 2 篇

2510.14032 2025-10-17 cs.CV 89%

Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding

Xiaoqian Shen, Wenxuan Zhang, Jun Chen, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology(卡布斯大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);video language model(abstract);分类 cs.CV

Comments NeurIPS 2025 (Spotlight). Webpage at https://xiaoqian-shen.github.io/Vgent

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13891 2025-10-17 cs.LG cs.AI 86%

K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding

Yifeng Yao, Yike Yun, Jing Wang, Huishuai Zhang, Dongyan Zhao, Ke Tian, Zhihao Wang, Minghui Qiu, Tao Wang

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) Bytedance(字节跳动)

专题命中 视频理解 :video understanding(title,abstract);long video(title)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 3 篇

2510.07670 2025-10-17 cs.CV cs.AI 57%

Ctrl-VI: Controllable Video Synthesis via Variational Inference

Haoyi Duan, Yunzhi Zhang, Yilun Du, Jiajun Wu

机构 * Stanford University(斯坦福大学) Harvard University(哈佛大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Project page: https://video-synthesis-variational.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14648 2025-10-17 cs.CV cs.AI 57%

In-Context Learning with Unpaired Clips for Instruction-based Video Editing

Xinyao Liao, Xianfang Zeng, Ziye Song, Zhoujie Fu, Gang Yu, Guosheng Lin

机构 * Nanyang Technological University(南洋理工大学) StepFun

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10623 2025-10-17 cs.LG cs.CV 57%

Flows and Diffusions on the Neural Manifold

Daniel Saragih, Deyu Cao, Tejas Balaji

机构 * Queen’s University and Vector Institute(女王大学和向量研究所) University of Tokyo(东京大学) University of Toronto(多伦多大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments 43 pages, 11 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 6 篇

2510.14179 2025-10-17 cs.CV cs.AI 83%

Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures

Yuancheng Xu, Wenqi Xian, Li Ma, Julien Philip, Ahmet Levent Taşel, Yiwei Zhao, Ryan Burgert, Mingming He, Oliver Hermann, Oliver Pilarski, Rahul Garg, Paul Debevec, Ning Yu

机构 * Eyeline Labs(Eyeline实验室)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Accepted to SIGGRAPH Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07944 2025-10-17 cs.CV 83%

CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving

Tianrui Zhang, Yichen Liu, Zilin Guo, Yuxin Guo, Jingcheng Ni, Chenjing Ding, Dan Xu, Lewei Lu, Zehuan Wu

机构 * Sensetime Research(商汤科技研究院) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09789 2025-10-17 cs.CV cs.AI 83%

On Equivariance and Fast Sampling in Video Diffusion Models Trained with Warped Noise

Chao Liu, Arash Vahdat

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18033 2025-10-17 cs.CV 79%

OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models

Dvir Samuel, Matan Levy, Nir Darshan, Gal Chechik, Rami Ben-Ari

机构 * Bar-Ilan University \& OriginAI Israel The Hebrew University of Jerusalem Israel Bar-Ilan University \& NVIDIA Research Israel Bar-Ilan University \& OriginAI The Hebrew University of Jerusalem Bar-Ilan University \& NVIDIA Research

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

Comments Accepted to SIGGRAPH ASIA 2025. Project Page: https://dvirsamuel.github.io/omnimattezero.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09656 2025-10-17 cs.CV 74%

KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation

Xingrui Wang, Jiang Liu, Ze Wang, Xiaodong Yu, Jialian Wu, Ximeng Sun, Yusheng Su, Alan Yuille, Zicheng Liu, Emad Barsoum

专题命中 视频扩散模型 :video diffusion(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23402 2025-10-17 cs.CV 57%

WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving

Ziyue Zhu, Zhanqian Wu, Zhenxin Zhu, Lijun Zhou, Haiyang Sun, Bing Wan, Kun Ma, Guang Chen, Hangjun Ye, Jin Xie, jian Yang

机构 * Nankai University(南开大学) Nanjing University, Suzhou(南京大学苏州校区)

专题命中 视频扩散模型 :video diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频问答 1 篇

2510.14672 2025-10-17 cs.CV 57%

VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning

Jinglei Zhang, Yuanfan Guo, Rolandos Alexandros Potamias, Jiankang Deng, Hang Xu, Chao Ma

机构 * Shanghai Jiao Tong University(上海交通大学) Noah’s Ark Lab(诺亚实验室) Imperial College London(伦敦帝国理工学院)

专题命中 视频问答 :video understanding(abstract);分类 cs.CV

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 长视频与时序推理 1 篇

2510.14624 2025-10-17 cs.CV 77%

Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference

Natan Bagrov, Eugene Khvedchenia, Borys Tymchenko, Shay Aharon, Lior Kadoch, Tomer Keren, Ofri Masad, Yonatan Geifman, Ran Zilberstein, Tuomas Rintamaki, Matthieu Le, Andrew Tao

机构 * NVIDIA

专题命中 长视频与时序推理 :video reasoning(abstract);video-language(abstract);long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏