arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6808 信号源:cs.CV, eess.IV, cs.MM

1. 视频扩散模型 1308 篇

2508.01698 2025-08-05 cs.CV 83%

Versatile Transition Generation with Image-to-Video Diffusion

Zuhao Yang, Jiahui Zhang, Yingchen Yu, Shijian Lu, Song Bai

机构 * Nanyang Technological University(南洋理工大学) ByteDance Inc.(字节跳动公司)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22360 2025-07-31 cs.CV cs.AI 83%

GVD: Guiding Video Diffusion Model for Scalable Video Distillation

Kunyang Li, Jeffrey A Chan Santiago, Sarinda Dhanesh Samarasinghe, Gaowen Liu, Mubarak Shah

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) Cisco Research(思科研究)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07982 2025-07-11 cs.CV cs.AI 83%

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Haoyu Wu, Diankun Wu, Tianyu He, Junliang Guo, Yang Ye, Yueqi Duan, Jiang Bian

机构 * Microsoft Research(微软研究院) Tsinghua University(清华大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments 18 pages, project page: https://GeometryForcing.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02860 2025-07-04 cs.CV 83%

Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching

Xin Zhou, Dingkang Liang, Kaijin Chen, Tianrui Feng, Xiwu Chen, Hongkai Lin, Yikang Ding, Feiyang Tan, Hengshuang Zhao, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) MEGVII Technology(美科七科技) University of Hong Kong(香港大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments The code is made available at https://github.com/H-EmbodVis/EasyCache. Project page: https://h-embodvis.github.io/EasyCache/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23858 2025-07-01 cs.CV 83%

VMoBA: Mixture-of-Block Attention for Video Diffusion Models

Jianzong Wu, Liang Hou, Haotian Yang, Xin Tao, Ye Tian, Pengfei Wan, Di Zhang, Yunhai Tong

机构 * Peking University(北京大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Code is at https://github.com/KwaiVGI/VMoBA

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17220 2025-06-24 cs.CV 83%

Emergent Temporal Correspondences from Video Diffusion Transformers

Jisu Nam, Soowon Son, Dahyun Chung, Jiyoung Kim, Siyoon Jin, Junhwa Hur, Seungryong Kim

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Project page is available at https://cvlab-kaist.github.io/DiffTrack

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15077 2025-06-18 cs.CV cs.AI 83%

Hardware-Friendly Static Quantization Method for Video Diffusion Transformers

Sanghyun Yi, Qingfeng Liu, Mostafa El-Khamy

机构 * Division of the Humanities and Social Sciences(人文与社会科学系) California Institute of Technology(加州理工学院) Device Solutions Research America(设备解决方案研究美国) Samsung Semiconductor Inc.(三星半导体公司)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Accepted to MIPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07280 2025-06-11 cs.CV cs.AI 83%

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models

Pablo Acuaviva, Aram Davtyan, Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, Alexandre Alahi, Paolo Favaro

机构 * University of Bern(伯恩大学) EPFL(瑞士联邦理工学院)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments 27 pages, 23 figures, 9 tables. Project page: https://pabloacuaviva.github.io/Gen2Gen/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06023 2025-06-09 cs.CV 83%

Restereo: Diffusion stereo video generation and restoration

Xingchang Huang, Ashish Kumar Singh, Florian Dubost, Cristina Nader Vasconcelos, Sakar Khattar, Liang Shi, Christian Theobalt, Cengiz Oztireli, Gurprit Singh

机构 * Max Planck Institute for Informatics(马克斯·普朗克信息研究所) VIA-Center Saarbücken(萨尔布吕肯VIA中心) University of Cambridge(剑桥大学) Google(谷歌) Google DeepMind(谷歌DeepMind)

专题命中 视频扩散模型 :video generation(title,abstract);video diffusion(abstract);分类 cs.CV

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04213 2025-06-06 cs.CV 83%

FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers

Xuanhua He, Quande Liu, Zixuan Ye, Weicai Ye, Qiulin Wang, Xintao Wang, Qifeng Chen, Pengfei Wan, Di Zhang, Kun Gai

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Kuaishou Technology(快手科技)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00830 2025-06-03 cs.CV 83%

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

Zhengcong Fei, Hao Jiang, Di Qiu, Baoxuan Gu, Youqiang Zhang, Jiahua Wang, Jialin Bai, Debang Li, Mingyuan Fan, Guibin Chen, Yahui Zhou

机构 * Kunlun Inc.(昆仑公司)

专题命中 视频扩散模型 :video diffusion(title,abstract);long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19151 2025-05-27 cs.GR cs.AI cs.CV 83%

SRDiffusion: Accelerate Video Diffusion Inference via Sketching-Rendering Cooperation

Shenggan Cheng, Yuanxin Wei, Lansong Diao, Yong Liu, Bujiao Chen, Lianghua Huang, Yu Liu, Wenyuan Yu, Jiangsu Du, Wei Lin, Yang You

机构 * National University of Singapore(国立新加坡大学) Sun Yat-sen University(中山大学) Alibaba Group(阿里巴巴集团)

专题命中 视频扩散模型 :video diffusion(title);video generation(abstract);text-to-video(abstract);分类 cs.CV

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14708 2025-05-22 cs.CV cs.AI 83%

DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance

Xuan Shen, Chenxia Han, Yufa Zhou, Yanyue Xie, Yifan Gong, Quanyi Wang, Yiwei Wang, Yanzhi Wang, Pu Zhao, Jiuxiang Gu

机构 * Northeastern University(东北大学) CUHK(香港中文大学) Duke University(杜克大学) Adobe Research(Adobe研究) NUIST(南京信息工程大学) UCM(墨西哥国立自治大学)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Preprint Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16375 2025-05-22 cs.CV 83%

Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing

Kaifeng Gao, Jiaxin Shi, Hanwang Zhang, Chunping Wang, Jun Xiao, Long Chen

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Accepted by ICML 2025. Code is available: https://github.com/Dawn-LX/CausalCache-VDM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21650 2025-05-14 cs.CV 83%

HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation

Haiyang Zhou, Wangbo Yu, Jiawen Guan, Xinhua Cheng, Yonghong Tian, Li Yuan

机构 * School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院) Peng Cheng Laboratory(鹏城实验室) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Project Homepage: https://zhouhyocean.github.io/holotime/ Code: https://github.com/PKU-YuanGroup/HoloTime

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10821 2025-05-06 cs.CV 83%

Tex4D: Zero-shot 4D Scene Texturing with Video Diffusion Models

Jingzhi Bao, Xueting Li, Ming-Hsuan Yang

机构 * CUHK-Shenzhen(香港中文大学(深圳)) NVIDIA(英伟达) UC Merced(加州大学默塞德分校)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Project page: https://tex4d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03048 2025-05-02 cs.CV 83%

Latte: Latent Diffusion Transformer for Video Generation

Xin Ma, Yaohui Wang, Xinyuan Chen, Gengyun Jia, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, Yu Qiao

机构 * Department of Data Science & AI, Faculty of Information Technology, Monash University(数据科学与人工智能系,信息科技学院,莫纳什大学) Shanghai AI Laboratory(上海人工智能实验室) Nanjing University of Posts and Telecommunications(南京邮电大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 视频扩散模型 :video generation(title,abstract);text-to-video(abstract);分类 cs.CV

Comments Accepted by Transactions on Machine Learning Research 2025; Project Page: https://maxin-cn.github.io/latte_project

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21669 2025-04-28 cs.CV 83%

Investigating Memorization in Video Diffusion Models

Chen Chen, Enhuai Liu, Daochang Liu, Mubarak Shah, Chang Xu

机构 * School of Computer Science, Faculty of Engineering, The University of Sydney, Australia(悉尼大学计算机科学学院,工程学院) School of Physics, Mathematics and Computing, The University of Western Australia, Australia(西澳大学物理、数学与计算学院) Center for Research in Computer Vision, University of Central Florida, USA(佛罗里达中央大学计算机视觉研究中心)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Accepted at DATA-FM Workshop @ ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04010 2025-04-08 cs.CV cs.LG 83%

DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion

Maksim Siniukov, Di Chang, Minh Tran, Hongkun Gong, Ashutosh Chaubey, Mohammad Soleymani

专题命中 视频扩散模型 :video generation(title,abstract);video diffusion(abstract);分类 cs.CV

Comments Project page: https://havent-invented.github.io/DiTaiListener

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02542 2025-04-08 cs.CV 83%

Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation

Fa-Ting Hong, Zunnan Xu, Zixiang Zhou, Jun Zhou, Xiu Li, Qin Lin, Qinglin Lu, Dan Xu

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02764 2025-04-04 cs.CV cs.AI 83%

Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model

Shengjun Zhang, Jinzhao Li, Xin Fei, Hao Liu, Yueqi Duan

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19462 2025-03-26 cs.CV 83%

AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset

Haiyu Zhang, Xinyuan Chen, Yaohui Wang, Xihui Liu, Yunhong Wang, Yu Qiao

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Project Page: https://aejion.github.io/accvideo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08244 2025-03-26 cs.CV 83%

FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis

Wonjoon Jin, Qi Dai, Chong Luo, Seung-Hwan Baek, Sunghyun Cho

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Our paper has been accepted to CVPR 2025. Website: https://jinwonjoon.github.io/flovd_site/ Code: https://github.com/JinWonjoon/FloVD

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19355 2025-03-13 cs.CV 83%

FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality

Zhengyao Lv, Chenyang Si, Junhao Song, Zhenyu Yang, Yu Qiao, Ziwei Liu, Kwan-Yee K. Wong

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14801 2025-03-05 cs.CV 83%

AVD2: Accident Video Diffusion for Accident Video Description

Cheng Li, Keyuan Zhou, Tong Liu, Yu Wang, Mingqiao Zhuang, Huan-ang Gao, Bu Jin, Hao Zhao

专题命中 视频扩散模型 :video diffusion(title,abstract);video understanding(abstract);分类 cs.CV

Comments ICRA 2025, Project Page: https://an-answer-tree.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06527 2025-02-21 cs.CV 83%

CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers

D. She, Mushui Liu, Jingxuan Pang, Jin Wang, Zhen Yang, Wanggui He, Guanghao Zhang, Yi Wang, Qihan Huang, Haobin Tang, Yunlong Yu, Siming Fu

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Section 4 in CustomVideoX Entity Region-Aware Enhancement has description errors. The compared methods data of Table I lacks other metrics

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17042 2025-02-18 cs.CV 83%

Adapting Image-to-Video Diffusion Models for Large-Motion Frame Interpolation

Luoxu Jin, Hiroshi Watanabe

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13301 2025-01-31 cs.CV cs.RO 83%

ARDuP: Active Region Video Diffusion for Universal Policies

Shuaiyi Huang, Mara Levy, Zhenyu Jiang, Anima Anandkumar, Yuke Zhu, Linxi Fan, De-An Huang, Abhinav Shrivastava

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Accepted by IROS 2024 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08225 2025-01-15 cs.CV 83%

FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors

Yabo Zhang, Xinpeng Zhou, Yihan Zeng, Hang Xu, Hui Li, Wangmeng Zuo

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments Code: https://github.com/YBYBZhang/FramePainter

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10150 2025-01-09 cs.CV 83%

Motion-Zero: Zero-Shot Moving Object Control Framework for Diffusion-Based Video Generation

Changgu Chen, Junwei Shu, Gaoqi He, Changbo Wang, Yang Li

专题命中 视频扩散模型 :video generation(title);video diffusion(abstract);text-to-video(abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏