arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-08-07 至 2025-08-07 共收录 8 信号源:cs.CV, eess.IV, cs.MM

1. 视频扩散模型 4 篇

2508.04467 2025-08-07 cs.CV 79%

4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation

Shuzhou Yang, Xiaodong Cun, Xiaoyu Li, Yaowei Li, Jian Zhang

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03123 2025-08-07 cs.CV 79%

Dual-Expert Consistency Model for Efficient and High-Quality Video Generation

Zhengyao Lv, Chenyang Si, Tianlin Pan, Zhaoxi Chen, Kwan-Yee K. Wong, Yu Qiao, Ziwei Liu

机构 * Nanjing University(南京大学) The University of Hong Kong(香港大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Chinese Academy of Sciences(中国科学院大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 视频扩散模型 :video generation(title);video diffusion(abstract);分类 cs.CV

Comments This paper has been accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08331 2025-08-07 cs.CV 79%

Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise

Ryan Burgert, Yuancheng Xu, Wenqi Xian, Oliver Pilarski, Pascal Clausen, Mingming He, Li Ma, Yitong Deng, Lingxiao Li, Mohsen Mousavi, Michael Ryoo, Paul Debevec, Ning Yu

机构 * Netflix Eyeline Studios(NetflixEyeline Studios) Netflix Stony Brook University(斯通布罗克大学) University of Maryland(马里兰大学) Stanford University(斯坦福大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);分类 cs.CV

Comments Accepted to CVPR'25 as Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04147 2025-08-07 cs.CV 74%

IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control

Lijuan Liu, Wenfa Li, Dongbo Zhang, Shuo Wang, Shaohui Jiao

专题命中 视频扩散模型 :video diffusion(title);分类 cs.CV

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频问答 1 篇

2508.04197 2025-08-07 cs.CV cs.AI 57%

Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective

Yan Zhang, Gangyan Zeng, Daiqing Wu, Huawen Shen, Binbin Li, Yu Zhou, Can Ma, Xiaojun Bi

机构 * Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences Beijing China School of Cyber Science Engineering, Nanjing University of Science VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Tianjin China Key Laboratory of Ethnic Language Intelligent Analysis Security Governance of MOE, Minzu University of China Beijing China Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences VCIP \& TMCC \& DISSec, College of Computer Science, Nankai University Security Governance of MOE, Minzu University of China

专题命中 视频问答 :video-language(abstract);分类 cs.CV

Comments Accepted by 2025 ACM MM

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 动作与事件理解 2 篇

2508.04049 2025-08-07 cs.CV 57%

Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation

Jiayi He, Xu Wang, Shengeng Tang, Yaxiong Wang, Lechao Cheng, Dan Guo

专题命中 动作与事件理解 :video generation(abstract);分类 cs.CV

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01304 2025-08-07 cs.CV 57%

Deep learning for action spotting in association football videos

Silvio Giancola, Anthony Cioppa, Bernard Ghanem, Marc Van Droogenbroeck

机构 * 1 Center of Excellence for Generative AI, IVUL, KAUST, Saudi Arabia 2 Montefiore Institute, Open-SportsLab, University of Li \`e ge, Belgium

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV

Comments 31 pages, 2 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频数据与评测 1 篇

2508.03955 2025-08-07 cs.CV 57%

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm

Lin Zhang, Zefan Cai, Yufan Zhou, Shentong Mo, Jinhong Lin, Cheng-En Wu, Yibing Wei, Yijing Zhang, Ruiyi Zhang, Wen Xiao, Tong Sun, Junjie Hu, Pedro Morgado

机构 * University of Wisconsin Madison(威斯康星大学麦迪逊分校) Carnegie Mellon University(卡内基梅隆大学) Luma AI Adobe Research(Adobe研究) Microsoft(微软)

专题命中 视频数据与评测 :text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏