arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-10 至 2025-10-10 共收录 6 信号源:cs.CV, eess.IV, cs.MM

1. 视频生成 6 篇

2510.05096 2025-10-10 cs.CV cs.AI cs.CL cs.MA cs.MM 81%

Paper2Video: Automatic Video Generation from Scientific Papers

Zeyu Zhu, Kevin Qinghong Lin, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(展示实验室,新加坡国立大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV、cs.MM

Comments Project Page: https://showlab.github.io/Paper2Video/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08527 2025-10-10 cs.CV 79%

FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control

Zhiyuan Zhang, Can Wang, Dongdong Chen, Jing Liao

机构 * City University of Hong Kong(香港城市大学) The University of Hong Kong(香港大学) Microsoft GenAI(微软生成人工智能)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

Comments Project Page: https://bestzzhang.github.io/FlexTraj

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23742 2025-10-10 cs.CV cs.AI 79%

MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement

Yufan Deng, Yuanyang Yin, Xun Guo, Yizhi Wang, Jacob Zhiyuan Fang, Shenghai Yuan, Yiding Yang, Angtian Wang, Bo Liu, Haibin Huang, Chongyang Ma

机构 * Intelligent Creation ByteDance(智能创作字节跳动)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

Comments Code: https://github.com/MAGREF-Video/MAGREF/; Project website: https://magref-video.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08143 2025-10-10 cs.CV 77%

UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution

Shian Du, Menghan Xia, Chang Liu, Quande Liu, Xintao Wang, Pengfei Wan, Xiangyang Ji

机构 * Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 视频生成 :video generation(abstract);video diffusion(abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12510 2025-10-10 cs.CV cs.MM 62%

PRVR: Partially Relevant Video Retrieval

Xianke Chen, Daizong Liu, Xun Yang, Xirong Li, Jianfeng Dong, Meng Wang, Xun Wang

机构 * School of Computer Science and Technology(计算机科学与技术学院) School of Statistics and Mathematics(统计学与数学学院) Zhejiang Gongshang University(浙江工商大学) Zhejiang Key Laboratory of Big Data and Future E-Commerce Technology(大数据与未来电子商务技术重点实验室) Wangxuan Institute of Computer Technology(王萱计算机技术研究所) School of Information Science and Technology(信息科学与技术学院) University of Science and Technology of China(中国科学技术大学) School of Information(信息学院) Renmin University of China(中国人民大学) School of Computer Science and Information Engineering(计算机科学与信息工程学院)

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV、cs.MM

Comments Accepted by TPAMI. The paper's homepage is https://github.com/HuiGuanLab/ms-sl-pp

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08568 2025-10-10 cs.RO cs.AI cs.CV 57%

NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos

Hongyu Li, Lingfeng Sun, Yafei Hu, Duy Ta, Jennifer Barry, George Konidaris, Jiahui Fu

机构 * Robotics and AI Institute(机器人与人工智能研究所) Brown University(布朗大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏