arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 2171 信号源:cs.CV, eess.IV, cs.MM

1. 视频生成 2171 篇

2104.11931 2021-04-27 cs.CV 57%

Adaptive Appearance Rendering

Mengyao Zhai, Ruizhi Deng, Jiacheng Chen, Lei Chen, Zhiwei Deng, Greg Mori

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted to BMVC 2018. arXiv admin note: substantial text overlap with arXiv:1712.01955

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.08849 2021-04-16 cs.CV cs.CL 57%

Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models

Po-Yao Huang, Mandela Patrick, Junjie Hu, Graham Neubig, Florian Metze, Alexander Hauptmann

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV

Comments accepted by NAACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.11588 2021-03-31 cs.MM cs.IR 57%

Generating Persuasive Visual Storylines for Promotional Videos

Chang Liu, Yi Dong, Han Yu, Zhiqi Shen, Zhanning Gao, Pan Wang, Changgong Zhang, Peiran Ren, Xuansong Xie, Lizhen Cui, Chunyan Miao

专题命中 视频生成 :video generation(abstract);分类 cs.MM

Comments 10 pages, accepted by The 28th ACM International Conference on Information and Knowledge Management (CIKM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.13329 2021-03-01 cs.CV 57%

Dual-MTGAN: Stochastic and Deterministic Motion Transfer for Image-to-Video Synthesis

Fu-En Yang, Jing-Cheng Chang, Yuan-Hao Lee, Yu-Chiang Frank Wang

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted to ICPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.02824 2021-01-15 cs.CV 57%

Support-set bottlenecks for video-text representation learning

Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander Hauptmann, João Henriques, Andrea Vedaldi

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV

Comments Accepted as spotlight paper at the International Conference on Learning Representations (ICLR) 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.09846 2020-11-30 cs.CV cs.CL cs.LG 57%

Everybody Sign Now: Translating Spoken Language to Photo Realistic Sign Language Video

Ben Saunders, Necati Cihan Camgoz, Richard Bowden

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.10727 2020-11-24 cs.CV 57%

Stochastic Talking Face Generation Using Latent Distribution Matching

Ravindra Yadav, Ashish Sardana, Vinay P Namboodiri, Rajesh M Hegde

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments InterSpeech 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.04208 2020-10-26 cs.CV 57%

Condensed Movies: Story Based Retrieval with Contextual Embeddings

Max Bain, Arsha Nagrani, Andrew Brown, Andrew Zisserman

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV

Comments Appears in: Asian Conference on Computer Vision 2020 (ACCV 2020) - Oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.12226 2020-10-23 cs.CV cs.LG 57%

Hierarchical Patch VAE-GAN: Generating Diverse Videos from a Single Sample

Shir Gur, Sagie Benaim, Lior Wolf

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.07634 2020-08-27 cs.CV 57%

DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat Rhythms

Hua Qi, Qing Guo, Felix Juefei-Xu, Xiaofei Xie, Lei Ma, Wei Feng, Yang Liu, Jianjun Zhao

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments 11 pages, 7 figures; This paper has been accepted to ACM-MM 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.11363 2020-08-27 cs.CV cs.LG 57%

How Do the Hearts of Deep Fakes Beat? Deep Fake Source Detection via Interpreting Residuals with Biological Signals

Umur Aybars Ciftci, Ilke Demir, Lijun Yin

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments To be published in the proceedings of 2020 IEEE/IAPR International Joint Conference on Biometrics (IJCB)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.11933 2020-06-23 cs.CV 57%

Lyric Video Analysis Using Text Detection and Tracking

Shota Sakaguchi, Jun Kato, Masataka Goto, Seiichi Uchida

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments 15 pages, 8 figures, DAS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.11487 2020-05-26 cs.CV 57%

Self-Training for Domain Adaptive Scene Text Detection

Yudi Chen, Wei Wang, Yu Zhou, Fei Yang, Dongbao Yang, Weiping Wang

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.07514 2020-04-17 cs.CV 57%

Local-Global Video-Text Interactions for Temporal Grounding

Jonghwan Mun, Minsu Cho, Bohyung Han

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV

Comments CVPR 2020; code available in https://github.com/JonghwanMun/LGI4temporalgrounding

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.02516 2020-01-17 cs.CV 57%

Learning a Text-Video Embedding from Incomplete and Heterogeneous Data

Antoine Miech, Ivan Laptev, Josef Sivic

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV

Comments The paper had a major update in January 2020 after a bug we found in the codebase that affected many results

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.04786 2019-07-29 cs.CV 57%

Talking Face Generation by Conditional Recurrent Adversarial Network

Yang Song, Jingwen Zhu, Dawei Li, Xiaolong Wang, Hairong Qi

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Project Page:http://web.eecs.utk.edu/~ysong18/projects/talkingface/talkingface.html

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.08845 2019-07-23 cs.CV 57%

Order Matters: Shuffling Sequence Generation for Video Prediction

Junyan Wang, Bingzhang Hu, Yang Long, Yu Guan

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments This manuscript has been accepted at BMVC 2019. See the project at https://github.com/andrewjywang/SEENet

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.02784 2019-04-19 cs.CV 57%

StoryGAN: A Sequential Conditional GAN for Story Visualization

Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, Jianfeng Gao

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.02631 2019-04-17 cs.CV stat.ML 57%

Sliced Wasserstein Generative Models

Jiqing Wu, Zhiwu Huang, Dinesh Acharya, Wen Li, Janine Thoma, Danda Pani Paudel, Luc Van Gool

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments This paper is accepted by CVPR 2019, accidentally uploaded as a new submission (arXiv:1904.05408, which has been withdrawn). The code is available at this https URL https:// github.com/musikisomorphie/swd.git

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.05408 2019-04-16 cs.CV 57%

Sliced Wasserstein Generative Models

Jiqing Wu, Zhiwu Huang, Dinesh Acharya, Wen Li, Janine Thoma, Danda Pani Paudel, Luc Van Gool

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments This paper is submitted to arxiv twice, thus withdraw one of the versions. See arXiv:1706.02631 instead

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.01890 2018-09-07 cs.CV cs.GR cs.LG stat.ML 57%

Full-body High-resolution Anime Generation with Progressive Structure-conditional Generative Adversarial Networks

Koichi Hamada, Kentaro Tachibana, Tianqi Li, Hiroto Honda, Yusuke Uchida

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted to ECCV 2018 Workshop: Computer Vision for Fashion, Art and Design. Project page is at https://dena.com/intl/anime-generation

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.02992 2018-08-10 cs.CV 57%

Controllable Image-to-Video Translation: A Case Study on Facial Expression Generation

Lijie Fan, Wenbing Huang, Chuang Gan, Junzhou Huang, Boqing Gong

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.01873 2018-03-29 cs.CV 57%

Every Smile is Unique: Landmark-Guided Diverse Smile Generation

Wei Wang, Xavier Alameda-Pineda, Dan Xu, Pascal Fua, Elisa Ricci, Nicu Sebe

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted as a poster in Conference on Computer Vision and Pattern Recognition (CVPR), 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.03664 2017-03-13 cs.CV cs.NE 57%

Parallel Multiscale Autoregressive Density Estimation

Scott Reed, Aäron van den Oord, Nal Kalchbrenner, Sergio Gómez Colmenarejo, Ziyu Wang, Dan Belov, Nando de Freitas

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.04809 2016-12-15 cs.CV 57%

Spectral video construction from RGB video: Application to Image Guided Neurosurgery

Md. Abul Hasnat, Jussi Parkkinen, Markku Hauta-Kasari

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Experiments were conducted in 2011, Paper rewritten with recent review in 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.09444 2016-12-07 cs.CV cs.AI cs.LG 57%

Contextual RNN-GANs for Abstract Reasoning Diagram Generation

Arnab Ghosh, Viveka Kulharia, Amitabha Mukerjee, Vinay Namboodiri, Mohit Bansal

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments To Appear in AAAI-17 and NIPS Workshop on Adversarial Training

详情

展开后加载摘要…

URL PDF HTML 收藏
1604.07939 2016-07-13 cs.MM cs.DB cs.IR 57%

Large-Scale Query-by-Image Video Retrieval Using Bloom Filters

Andre Araujo, Jason Chaves, Haricharan Lakshman, Roland Angst, Bernd Girod

专题命中 视频生成 :long video(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07763 2026-07-10 cs.LG 新提交 56%

Unlocking Temporal Generalization in Hamiltonian Video Dynamics Models

解锁哈密顿视频动力学模型中的时间泛化

Eli Laird, Corey Clark

机构 * Department of Computer Science, Southern Methodist University(南卫理公会大学计算机科学系)

专题命中 视频生成 :video generation(abstract,comments)

AI总结 研究世界模型在可变时间分辨率下预测动力学的问题,利用哈密顿生成网络(HGN),指出其在非保守环境中时间泛化失效的问题及原因,通过针对性修复实现稳定动力学预测,推荐连续时间视频生成中时间泛化的策略。

Comments To appear in the 1st Workshop on Physics-Aware Video Generation and Restoration at the 28th International Conference on Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.26581 2026-08-28 cs.LG 新提交 50%

Activation Outliers Matter: Robust Recovery for Quantized Multimodal LLMs

激活异常值很重要:量化多模态大语言模型的鲁棒恢复

Tanzila Rahman, Mehran Taghian Jazi, Yunke Peng, Zhuang Ma, Anandharaju Durai Raju, Yao Wang, Xing Huang, Hei Yi Mak, Shadan Golestan, Hoang Le, Yonghan Dong, Wei Guo, Yaoyuan Wang

机构 * Huawei(华为)

专题命中 视频生成 :video generation(abstract)

AI总结 本研究针对多模态大语言模型超低比特量化的性能损失问题,提出Residual Fallback Quantization框架,可有效恢复MXFP4、HiF4量化下的性能,缩小与BF16基线的差距。

Comments 14 Pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.24329 2026-08-26 cs.HC 新提交 50%

How Do Professional Editors Evaluate the Editing Quality of AI-Generated Cinematic Video Ads?

专业编辑如何评估AI生成的影视视频广告的剪辑质量?

Po-Ming Law, Weizhi Li, Arpit Narechania

专题命中 视频生成 :video generation(abstract)

AI总结 该研究针对AI生成影视广告缺乏细粒度评估框架的问题,构建两阶段生成流程,通过专业编辑评价得出六个剪辑质量维度,为相关评估与生成提供指导。

详情

展开后加载摘要…

URL PDF HTML 收藏