arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-16 至 2025-10-16 共收录 8 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1 篇

2505.16836 2025-10-16 cs.CV cs.AI 57%

Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning

Fanrui Zhang, Dian Li, Qiang Zhang, Jun Chen, Gang Liu, Junxiong Lin, Jiahong Yan, Jiawei Liu, Zheng-Jun Zha

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, USTC(脑启发智能感知与认知国家重点实验室,中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) Tencent QQ(腾讯QQ) Fudan University(复旦大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 34 pages, 25 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 4 篇

2502.03207 2025-10-16 cs.CV cs.GR 83%

MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent

Xinyao Liao, Xianfang Zeng, Liao Wang, Gang Yu, Guosheng Lin, Chi Zhang

机构 * Nanyang Technological University(南洋理工大学) StepFun Westlake University(西湖大学)

专题命中 视频生成 :video generation(title,abstract);video diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02512 2025-10-16 cs.RO cs.CV eess.IV 81%

QuaDreamer: Controllable Panoramic Video Generation for Quadruped Robots

Sheng Wu, Fei Teng, Hao Shi, Qi Jiang, Kai Luo, Kaiwei Wang, Kailun Yang

机构 * Hunan University(湖南大学) Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV、eess.IV

Comments Accepted to CoRL 2025. The source code and model weights will be publicly available at https://github.com/losehu/QuaDreamer

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13809 2025-10-16 cs.CV 79%

PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning

Sihui Ji, Xi Chen, Xin Tao, Pengfei Wan, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

Comments Project Page: https://sihuiji.github.io/PhysMaster-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13652 2025-10-16 cs.CV 57%

EditCast3D: Single-Frame-Guided 3D Editing with Video Propagation and View Selection

Huaizhi Qu, Ruichen Zhang, Shuqing Luo, Luchao Qi, Zhihao Zhang, Xiaoming Liu, Roni Sengupta, Tianlong Chen

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Michigan State University(密歇根州立大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频扩散模型 2 篇

2504.12626 2025-10-16 cs.CV 83%

Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models

Lvmin Zhang, Shengqu Cai, Muyang Li, Gordon Wetzstein, Maneesh Agrawala

机构 * Stanford University(斯坦福大学) MIT(麻省理工学院)

专题命中 视频扩散模型 :video diffusion(title,abstract);video generation(abstract);分类 cs.CV

Comments https://github.com/lllyasviel/FramePack

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13678 2025-10-16 cs.CV 57%

FlashWorld: High-quality 3D Scene Generation within Seconds

Xinyang Li, Tengfei Wang, Zixiao Gu, Shengchuan Zhang, Chunchao Guo, Liujuan Cao

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学) Tencent(腾讯) Yes Lab, Fudan University(复旦大学Yes实验室)

专题命中 视频扩散模型 :video diffusion(abstract);分类 cs.CV

Comments Project Page: https://imlixinyang.github.io/FlashWorld-Project-Page/

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频数据与评测 1 篇

2510.13042 2025-10-16 cs.CV cs.AI 79%

SeqBench: Benchmarking Sequential Narrative Generation in Text-to-Video Models

Zhengxu Tang, Zizheng Wang, Luning Wang, Zitao Shuai, Chenhao Zhang, Siyu Qian, Yirui Wu, Bohao Wang, Haosong Rao, Zhenyu Yang, Chenwei Wu

机构 * University of Michigan(密歇根大学) Northeastern University(东北大学) University of Washington(华盛顿大学) Harvard University(哈佛大学) Beijing Jiaotong University(北京交通大学) University of Rochester(罗切斯特大学) School of Earth Sciences(地球科学学院)

专题命中 视频数据与评测 :text-to-video(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏