arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 2160 信号源:cs.CV, eess.IV, cs.MM

1. 视频生成 2160 篇

2508.14033 2025-08-20 cs.CV 74%

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing

Shaoshu Yang, Zhe Kong, Feng Gao, Meng Cheng, Xiangyu Liu, Yong Zhang, Zhuoliang Kang, Wenhan Luo, Xunliang Cai, Ran He, Xiaoming Wei

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Meituan New Laboratory of Pattern Recognition (NLPR), CASIA(美团模式识别新实验室(NLPR),中国科学院自动化所) Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Meituan Division of AMC and Department of ECE, HKUST(美团AMC部门和香港科技大学电子与计算机工程系) State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室,中国科学院自动化所) Meituan School of Artificial Intelligence, University of Chinese Academy of Sciences(美团人工智能学院,中国科学院大学) Meituan(美团)

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments 11 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17414 2025-07-14 cs.CV 74%

X-Dancer: Expressive Music to Human Dance Video Generation

Zeyuan Chen, Hongyi Xu, Guoxian Song, You Xie, Chenxu Zhang, Xin Chen, Chao Wang, Di Chang, Linjie Luo

机构 * UC San Diego(圣迭戈大学) ByteDance(字节跳动) University of Southern California(南加州大学)

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments ICCV 2025. Project Page: https://zeyuan-chen.com/X-Dancer/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04036 2025-07-08 cs.CV 74%

PresentAgent: Multimodal Agent for Presentation Video Generation

Jingwei Shi, Zeyu Zhang, Biao Wu, Yanjie Liang, Meng Fang, Ling Chen, Yang Zhao

机构 * AI Geeks, Australia Australian Artificial Intelligence Institute, Australia(澳大利亚人工智能研究所) University of Liverpool, United Kingdom(利物浦大学) La Trobe University, Australia(拉特罗布大学)

专题命中 视频生成 :video generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22046 2025-07-01 cs.CV 74%

LatentMove: Towards Complex Human Movement Video Generation

Ashkan Taghipour, Morteza Ghahremani, Mohammed Bennamoun, Farid Boussaid, Aref Miri Rekavandi, Zinuo Li, Qiuhong Ke, Hamid Laga

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments The authors are withdrawing this paper due to major issues in the experiments and methodology. To prevent citation of this outdated and flawed version, we have decided to remove it while we work on a substantial revision. Thank you

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04956 2025-06-06 cs.CV 74%

FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation

Huihan Wang, Zhiwen Yang, Hui Zhang, Dan Zhao, Bingzheng Wei, Yan Xu

机构 * School of Biological Science and Medical Engineering, State Key Laboratory of Software Development Environment, Key Laboratory of Biomechanics and Mechanobiology of Ministry of Education, Beijing Advanced Innovation Center for Biomedical Engineering, Beihang University, Beijing 100191, China(生物科学与医学工程学院、软件开发环境国家重点实验室、教育部生物力学与机械生物学重点实验室、北京生物医学工程先进创新中心、北京航空航天大学) Department of Biomedical Engineering, Tsinghua University, Beijing 100084, China(生物医学工程系、清华大学) Department of Gynecology Oncology, National Cancer Center/National Clinical Research Center for Cancer/Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing 100021, China(妇科肿瘤科、国家癌症中心/国家癌症临床研究中心/癌症医院、中国医学科学院和北京协和医学院) ByteDance Inc., Beijing 100098, China(字节跳动公司)

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments This paper has been early accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03140 2025-06-04 cs.CV 74%

CamCloneMaster: Enabling Reference-based Camera Control for Video Generation

Yawen Luo, Jianhong Bai, Xiaoyu Shi, Menghan Xia, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai, Tianfan Xue

机构 * The Chinese University of Hong Kong(香港中文大学) Zhejiang University(浙江大学) Kuaishou Technology(快手科技)

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments Project Page: https://camclonemaster.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01103 2025-06-03 cs.CV 74%

DeepVerse: 4D Autoregressive Video Generation as a World Model

Junyi Chen, Haoyi Zhu, Xianglong He, Yifan Wang, Jianjun Zhou, Wenzheng Chang, Yang Zhou, Zizun Li, Zhoujie Fu, Jiangmiao Pang, Tong He

机构 * SJTU(上海交通大学) Shanghai AI Lab(上海人工智能实验室) USTC(中国科学技术大学) THU(清华大学) ZJU(浙江大学) FDU(福建师范大学) NTU(国立台湾大学)

专题命中 视频生成 :video generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15880 2025-05-26 cs.CV 74%

Challenger: Affordable Adversarial Driving Video Generation

Zhiyuan Xu, Bohan Li, Huan-ang Gao, Mingju Gao, Yong Chen, Ming Liu, Chenxu Yan, Hang Zhao, Shuo Feng, Hao Zhao

机构 * AIR, Tsinghua(清华大学人工智能研究院) UCAS(中国科学技术大学) SJTU(上海交通大学) EIT, Ningbo(宁波工程学院) Geely Auto(吉利汽车) IIIS, Tsinghua(清华大学人工智能研究院) DA, Tsinghua(清华大学人工智能研究院)

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments Project page: https://pixtella.github.io/Challenger/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13211 2025-05-20 cs.CV cs.AI 74%

MAGI-1: Autoregressive Video Generation at Scale

Sand. ai, Hansi Teng, Hongyu Jia, Lei Sun, Lingzhi Li, Maolin Li, Mingqiu Tang, Shuai Han, Tianning Zhang, W. Q. Zhang, Weifeng Luo, Xiaoyang Kang, Yuchen Sun, Yue Cao, Yunpeng Huang, Yutong Lin, Yuxin Fang, Zewei Tao, Zheng Zhang, Zhongshu Wang, Zixun Liu, Dai Shi, Guoli Su, Hanwen Sun, Hong Pan, Jie Wang, Jiexin Sheng, Min Cui, Min Hu, Ming Yan, Shucheng Yin, Siran Zhang, Tingting Liu, Xianping Yin, Xiaoyu Yang, Xin Song, Xuan Hu, Yankai Zhang, Yuqiao Li

专题命中 视频生成 :video generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22357 2025-03-31 cs.CV 74%

EchoFlow: A Foundation Model for Cardiac Ultrasound Image and Video Generation

Hadrien Reynaud, Alberto Gomez, Paul Leeson, Qingjie Meng, Bernhard Kainz

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12761 2025-03-17 cs.CV cs.AI cs.LG 74%

SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation

Jaehong Yoon, Shoubin Yu, Vaidehi Patil, Huaxiu Yao, Mohit Bansal

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments ICLR 2025; The first two authors contributed equally; Project page: https://safree-safe-t2i-t2v.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09733 2025-03-14 cs.CV 74%

I2V3D: Controllable image-to-video generation with 3D guidance

Zhiyuan Zhang, Dongdong Chen, Jing Liao

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments Project page: https://bestzzhang.github.io/I2V3D

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08307 2025-03-13 cs.CV 74%

$^R$FLAV: Rolling Flow matching for infinite Audio Video generation

Alex Ergasti, Giuseppe Gabriele Tarollo, Filippo Botti, Tomaso Fontanini, Claudio Ferrari, Massimo Bertozzi, Andrea Prati

专题命中 视频生成 :video generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03751 2025-03-06 cs.CV cs.GR 74%

GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control

Xuanchi Ren, Tianchang Shen, Jiahui Huang, Huan Ling, Yifan Lu, Merlin Nimier-David, Thomas Müller, Alexander Keller, Sanja Fidler, Jun Gao

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments To appear in CVPR 2025. Website: https://research.nvidia.com/labs/toronto-ai/GEN3C/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00948 2025-03-04 cs.CV 74%

Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think

Jie Tian, Xiaoye Qu, Zhenyi Lu, Wei Wei, Sichen Liu, Yu Cheng

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments Accepted by CVPR2025

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01101 2025-02-18 cs.CV cs.AI 74%

VidSketch: Hand-drawn Sketch-Driven Video Generation with Diffusion Control

Lifan Jiang, Shuang Chen, Boxi Wu, Xiaotong Guan, Jiahui Zhang

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments 17pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10687 2025-01-22 cs.CV 74%

EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Linrui Tian, Siqi Hu, Qi Wang, Bang Zhang, Liefeng Bo

专题命中 视频生成 :video generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15220 2024-12-23 cs.MM cs.SD eess.AS 74%

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Haohe Liu, Gael Le Lan, Xinhao Mei, Zhaoheng Ni, Anurag Kumar, Varun Nagaraja, Wenwu Wang, Mark D. Plumbley, Yangyang Shi, Vikas Chandra

专题命中 视频生成 :video generation(title);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03520 2024-12-10 cs.CV 74%

Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention

Hannan Lu, Xiaohe Wu, Shudong Wang, Xiameng Qin, Xinyu Zhang, Junyu Han, Wangmeng Zuo, Ji Tao

专题命中 视频生成 :video generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12663 2024-11-20 cs.CV cs.AI cs.LG 74%

PoM: Efficient Image and Video Generation with the Polynomial Mixer

David Picard, Nicolas Dufour

专题命中 视频生成 :video generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16709 2024-09-26 cs.CV 74%

Pose-Guided Fine-Grained Sign Language Video Generation

Tongkai Shi, Lianyu Hu, Fanhua Shang, Jichao Feng, Peidong Liu, Wei Feng

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07605 2024-08-15 cs.CV 74%

Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving

Yuqing Wen, Yucheng Zhao, Yingfei Liu, Binyuan Huang, Fan Jia, Yanhui Wang, Chi Zhang, Tiancai Wang, Xiaoyan Sun, Xiangyu Zhang

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments Project page: https://panacea-ad.github.io/. arXiv admin note: text overlap with arXiv:2311.16813

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06035 2024-08-13 cs.CV cs.LG 74%

RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks

Partha Ghosh, Soubhik Sanyal, Cordelia Schmid, Bernhard Schölkopf

专题命中 视频生成 :video generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00808 2024-06-04 cs.CV 74%

EchoNet-Synthetic: Privacy-preserving Video Generation for Safe Medical Data Sharing

Hadrien Reynaud, Qingjie Meng, Mischa Dombrowski, Arijit Ghosh, Thomas Day, Alberto Gomez, Paul Leeson, Bernhard Kainz

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments Accepted at MICCAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16537 2024-05-28 cs.CV 74%

I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion Models

Wenqi Ouyang, Yi Dong, Lei Yang, Jianlou Si, Xingang Pan

专题命中 视频生成 :video diffusion(title);分类 cs.CV

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.04946 2024-04-09 cs.CV 74%

AnimateZoo: Zero-shot Video Generation of Cross-Species Animation via Subject Alignment

Yuanfeng Xu, Yuhao Chen, Zhongzhan Huang, Zijian He, Guangrun Wang, Philip Torr, Liang Lin

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments Technical report,15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11700 2024-03-25 cs.MM 74%

Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing

Juan Zhang, Jiahao Chen, Cheng Wang, Zhiwang Yu, Tangquan Qi, Can Liu, Di Wu

专题命中 视频生成 :video generation(title);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.00521 2024-02-13 cs.CV cs.AI cs.LG 74%

StyleLipSync: Style-based Personalized Lip-sync Video Generation

Taekyung Ki, Dongchan Min

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments International Conference on Computer Vision (ICCV) 2023. Project page: https://stylelipsync.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14294 2023-12-15 cs.CV 74%

Decouple Content and Motion for Conditional Image-to-Video Generation

Cuifeng Shen, Yulu Gan, Chen Chen, Xiongwei Zhu, Lele Cheng, Tingting Gao, Jinzhi Wang

专题命中 视频生成 :video generation(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.06225 2023-12-12 cs.CV cs.AI 74%

DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation

Fa-Ting Hong, Li Shen, Dan Xu

专题命中 视频生成 :video generation(title);分类 cs.CV

Comments Accepted at TPAMI; CVPR 2022 extension

详情

展开后加载摘要…

URL PDF HTML 收藏