arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6808 信号源:cs.CV, eess.IV, cs.MM

1. 视频生成 2156 篇

2011.09846 2020-11-30 cs.CV cs.CL cs.LG 57%

Everybody Sign Now: Translating Spoken Language to Photo Realistic Sign Language Video

Ben Saunders, Necati Cihan Camgoz, Richard Bowden

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.10727 2020-11-24 cs.CV 57%

Stochastic Talking Face Generation Using Latent Distribution Matching

Ravindra Yadav, Ashish Sardana, Vinay P Namboodiri, Rajesh M Hegde

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments InterSpeech 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.04208 2020-10-26 cs.CV 57%

Condensed Movies: Story Based Retrieval with Contextual Embeddings

Max Bain, Arsha Nagrani, Andrew Brown, Andrew Zisserman

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV

Comments Appears in: Asian Conference on Computer Vision 2020 (ACCV 2020) - Oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.12226 2020-10-23 cs.CV cs.LG 57%

Hierarchical Patch VAE-GAN: Generating Diverse Videos from a Single Sample

Shir Gur, Sagie Benaim, Lior Wolf

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.07634 2020-08-27 cs.CV 57%

DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat Rhythms

Hua Qi, Qing Guo, Felix Juefei-Xu, Xiaofei Xie, Lei Ma, Wei Feng, Yang Liu, Jianjun Zhao

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments 11 pages, 7 figures; This paper has been accepted to ACM-MM 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.11363 2020-08-27 cs.CV cs.LG 57%

How Do the Hearts of Deep Fakes Beat? Deep Fake Source Detection via Interpreting Residuals with Biological Signals

Umur Aybars Ciftci, Ilke Demir, Lijun Yin

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments To be published in the proceedings of 2020 IEEE/IAPR International Joint Conference on Biometrics (IJCB)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.11933 2020-06-23 cs.CV 57%

Lyric Video Analysis Using Text Detection and Tracking

Shota Sakaguchi, Jun Kato, Masataka Goto, Seiichi Uchida

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments 15 pages, 8 figures, DAS 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.11487 2020-05-26 cs.CV 57%

Self-Training for Domain Adaptive Scene Text Detection

Yudi Chen, Wei Wang, Yu Zhou, Fei Yang, Dongbao Yang, Weiping Wang

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.07514 2020-04-17 cs.CV 57%

Local-Global Video-Text Interactions for Temporal Grounding

Jonghwan Mun, Minsu Cho, Bohyung Han

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV

Comments CVPR 2020; code available in https://github.com/JonghwanMun/LGI4temporalgrounding

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.02516 2020-01-17 cs.CV 57%

Learning a Text-Video Embedding from Incomplete and Heterogeneous Data

Antoine Miech, Ivan Laptev, Josef Sivic

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV

Comments The paper had a major update in January 2020 after a bug we found in the codebase that affected many results

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.04786 2019-07-29 cs.CV 57%

Talking Face Generation by Conditional Recurrent Adversarial Network

Yang Song, Jingwen Zhu, Dawei Li, Xiaolong Wang, Hairong Qi

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Project Page:http://web.eecs.utk.edu/~ysong18/projects/talkingface/talkingface.html

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.08845 2019-07-23 cs.CV 57%

Order Matters: Shuffling Sequence Generation for Video Prediction

Junyan Wang, Bingzhang Hu, Yang Long, Yu Guan

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments This manuscript has been accepted at BMVC 2019. See the project at https://github.com/andrewjywang/SEENet

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.02784 2019-04-19 cs.CV 57%

StoryGAN: A Sequential Conditional GAN for Story Visualization

Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, Jianfeng Gao

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.02631 2019-04-17 cs.CV stat.ML 57%

Sliced Wasserstein Generative Models

Jiqing Wu, Zhiwu Huang, Dinesh Acharya, Wen Li, Janine Thoma, Danda Pani Paudel, Luc Van Gool

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments This paper is accepted by CVPR 2019, accidentally uploaded as a new submission (arXiv:1904.05408, which has been withdrawn). The code is available at this https URL https:// github.com/musikisomorphie/swd.git

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.05408 2019-04-16 cs.CV 57%

Sliced Wasserstein Generative Models

Jiqing Wu, Zhiwu Huang, Dinesh Acharya, Wen Li, Janine Thoma, Danda Pani Paudel, Luc Van Gool

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments This paper is submitted to arxiv twice, thus withdraw one of the versions. See arXiv:1706.02631 instead

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.01890 2018-09-07 cs.CV cs.GR cs.LG stat.ML 57%

Full-body High-resolution Anime Generation with Progressive Structure-conditional Generative Adversarial Networks

Koichi Hamada, Kentaro Tachibana, Tianqi Li, Hiroto Honda, Yusuke Uchida

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted to ECCV 2018 Workshop: Computer Vision for Fashion, Art and Design. Project page is at https://dena.com/intl/anime-generation

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.02992 2018-08-10 cs.CV 57%

Controllable Image-to-Video Translation: A Case Study on Facial Expression Generation

Lijie Fan, Wenbing Huang, Chuang Gan, Junzhou Huang, Boqing Gong

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.01873 2018-03-29 cs.CV 57%

Every Smile is Unique: Landmark-Guided Diverse Smile Generation

Wei Wang, Xavier Alameda-Pineda, Dan Xu, Pascal Fua, Elisa Ricci, Nicu Sebe

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Accepted as a poster in Conference on Computer Vision and Pattern Recognition (CVPR), 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.03664 2017-03-13 cs.CV cs.NE 57%

Parallel Multiscale Autoregressive Density Estimation

Scott Reed, Aäron van den Oord, Nal Kalchbrenner, Sergio Gómez Colmenarejo, Ziyu Wang, Dan Belov, Nando de Freitas

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.04809 2016-12-15 cs.CV 57%

Spectral video construction from RGB video: Application to Image Guided Neurosurgery

Md. Abul Hasnat, Jussi Parkkinen, Markku Hauta-Kasari

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments Experiments were conducted in 2011, Paper rewritten with recent review in 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.09444 2016-12-07 cs.CV cs.AI cs.LG 57%

Contextual RNN-GANs for Abstract Reasoning Diagram Generation

Arnab Ghosh, Viveka Kulharia, Amitabha Mukerjee, Vinay Namboodiri, Mohit Bansal

专题命中 视频生成 :video generation(abstract);分类 cs.CV

Comments To Appear in AAAI-17 and NIPS Workshop on Adversarial Training

详情

展开后加载摘要…

URL PDF HTML 收藏
1604.07939 2016-07-13 cs.MM cs.DB cs.IR 57%

Large-Scale Query-by-Image Video Retrieval Using Bloom Filters

Andre Araujo, Jason Chaves, Haricharan Lakshman, Roland Angst, Bernd Girod

专题命中 视频生成 :long video(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07763 2026-07-10 cs.LG 新提交 56%

Unlocking Temporal Generalization in Hamiltonian Video Dynamics Models

解锁哈密顿视频动力学模型中的时间泛化

Eli Laird, Corey Clark

机构 * Department of Computer Science, Southern Methodist University(南卫理公会大学计算机科学系)

专题命中 视频生成 :video generation(abstract,comments)

AI总结 研究世界模型在可变时间分辨率下预测动力学的问题,利用哈密顿生成网络(HGN),指出其在非保守环境中时间泛化失效的问题及原因,通过针对性修复实现稳定动力学预测,推荐连续时间视频生成中时间泛化的策略。

Comments To appear in the 1st Workshop on Physics-Aware Video Generation and Restoration at the 28th International Conference on Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14398 2026-08-21 cs.AI 版本更新 50%

ChronoAgentic: A Code-based Multi-Agent World Simulator for Physically Grounded Simulation Construction

编码智能体作为世界模拟器

Hongyu Wang, Jingquan Wang, Ashvin Anilkumar, Bocheng Zou, Radu Serban, Dan Negrut

机构 * Department of Mechanical & Aerospace Engineering, University of Wisconsin-Madison(威斯康星大学麦迪逊分校机械与航空航天工程系) School of Computer, Data, and Information Sciences, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机、数据与信息科学学院)

专题命中 视频生成 :text-to-video(abstract)

AI总结 提出一个通过可执行模拟代码构建基于物理的世界模型的智能体框架,协调规划、代码生成、视觉审查和物理分析智能体,迭代修正代码以满足物理约束,在物理准确性、指令忠实度和视觉质量上超越视频模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12453 2026-08-18 cs.LG 版本更新 50%

Time-Correlated Video Bridge Matching

时间相关视频桥匹配

Viacheslav Vasilev, Arseny Ivanov, Nikita Gushchin, Maria Kovaleva, Alexander Korotin

机构 * Kandinsky Lab(Kandinsky实验室) Applied AI Institute(应用人工智能研究所) HSE University(高等经济大学)

专题命中 视频生成 :video generation(abstract)

AI总结 本文提出时间相关视频桥匹配框架,解决视频生成中时间相关序列建模问题,通过建模序列间依赖性和时间相关性,提升视频生成质量与重建精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28203 2026-07-31 cs.HC 新提交 50%

Student Perceptions and Preferences Regarding AI-Generated Instructional Videos in Computing Education

计算教育中学生对AI生成教学视频的感知与偏好

Esse Ciego, Shubbhi Taneja, Wilson Wong, Amanpreet Kapoor

专题命中 视频生成 :video generation(abstract)

AI总结 本研究通过对170名计算专业学生的调查,探究其对Knowlify生成的Markdown教学AI视频的感知与偏好,发现学生认可视频质量但对其课堂广泛采用存顾虑,明确了AI视频的适用场景与潜在问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23969 2026-07-31 cs.RO 版本更新 50%

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

LeapBot-WA:通过预测性潜在对齐实现的世界锚定动作模型

Pei Liu, Nan Zheng, Lang Zhang, Daojie Peng, Yanan Zhang, Feilong Kong, Mingyue Feng, Jiachao Liu, Yaonong Wang, Qifeng Chen, Jun Ma

专题命中 视频生成 :video generation(abstract)

AI总结 研究针对世界动作模型依赖像素级视频生成的瓶颈,提出LeapBot-WA,通过预测性潜在对齐建立新范式,引入ISAE弥合模态差距,设计MoT架构,在多数据集上表现出色,实现高效强大的潜在中心范式及零样本鲁棒性和现实世界迁移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26903 2026-07-30 cs.AI cs.RO 新提交 50%

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

从被动视频到可编辑体验:具身智能的物理基础体验合成

Jia Luo

专题命中 视频生成 :video generation(abstract)

AI总结 针对具身AI的数据瓶颈,提出Pegasus低资源框架,通过结构化知识传递将人类操作视频转化为机器人可学习数据,经多基准与机器人评估验证其跨具身翻译及数据生成的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12995 2026-07-21 cs.RO 版本更新 50%

GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training

GenHOI: 通过模仿生成视频实现接触感知的人形机器人-物体交互,无需任务特定训练

Zhihai Bi, Qiang Zhang, Guoyang Zhao, Jiahang Cao, Xueyin Luo, Yushan Zhang, Jinglan Xu, Ruoyu Geng, Yulin Li, Andrew F. Luo, Jun Ma

机构 * The University of Tokyo(东京大学) National University of Singapore(新加坡国立大学) University of California, Los Angeles(加州大学洛杉矶分校) Tsinghua University(清华大学)

专题命中 视频生成 :video generation(abstract)

AI总结 提出GenHOI框架,通过模仿单个生成视频实现人形机器人零样本执行多种物体交互任务,无需任务特定训练或物理演示数据,利用接触事件和手-物接触区域编码为几何约束优化轨迹。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.09803 2026-07-14 cs.LG cs.CL 新提交 50%

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation

自回归生成中自我修正盲点的频谱起源

Ingrid Petrova, Luan Vejsiu

机构 * European University of Tirana(地拉那欧洲大学)

专题命中 视频生成 :video generation(abstract)

AI总结 研究大型自回归语言模型自我修正盲点问题,提出频谱代数理论SPARC,定义错误传播算子,推导激活阈值,证明基于强化学习的验证器-校正器训练收敛条件,实验验证定理,频谱预测与盲点率匹配。

详情

展开后加载摘要…

URL PDF HTML 收藏