arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-10 至 2025-10-10 共收录 13 信号源:cs.CV, eess.IV, cs.MM

1. 视频生成 6 篇

2510.05096 2025-10-10 cs.CV cs.AI cs.CL cs.MA cs.MM 81%

Paper2Video: Automatic Video Generation from Scientific Papers

Zeyu Zhu, Kevin Qinghong Lin, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(展示实验室,新加坡国立大学)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV、cs.MM

Comments Project Page: https://showlab.github.io/Paper2Video/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08527 2025-10-10 cs.CV 79%

FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control

Zhiyuan Zhang, Can Wang, Dongdong Chen, Jing Liao

机构 * City University of Hong Kong(香港城市大学) The University of Hong Kong(香港大学) Microsoft GenAI(微软生成人工智能)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

Comments Project Page: https://bestzzhang.github.io/FlexTraj

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23742 2025-10-10 cs.CV cs.AI 79%

MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement

Yufan Deng, Yuanyang Yin, Xun Guo, Yizhi Wang, Jacob Zhiyuan Fang, Shenghai Yuan, Yiding Yang, Angtian Wang, Bo Liu, Haibin Huang, Chongyang Ma

机构 * Intelligent Creation ByteDance(智能创作字节跳动)

专题命中 视频生成 :video generation(title,abstract);分类 cs.CV

Comments Code: https://github.com/MAGREF-Video/MAGREF/; Project website: https://magref-video.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08143 2025-10-10 cs.CV 77%

UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution

Shian Du, Menghan Xia, Chang Liu, Quande Liu, Xintao Wang, Pengfei Wan, Xiangyang Ji

机构 * Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 视频生成 :video generation(abstract);video diffusion(abstract);text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12510 2025-10-10 cs.CV cs.MM 62%

PRVR: Partially Relevant Video Retrieval

Xianke Chen, Daizong Liu, Xun Yang, Xirong Li, Jianfeng Dong, Meng Wang, Xun Wang

机构 * School of Computer Science and Technology(计算机科学与技术学院) School of Statistics and Mathematics(统计学与数学学院) Zhejiang Gongshang University(浙江工商大学) Zhejiang Key Laboratory of Big Data and Future E-Commerce Technology(大数据与未来电子商务技术重点实验室) Wangxuan Institute of Computer Technology(王萱计算机技术研究所) School of Information Science and Technology(信息科学与技术学院) University of Science and Technology of China(中国科学技术大学) School of Information(信息学院) Renmin University of China(中国人民大学) School of Computer Science and Information Engineering(计算机科学与信息工程学院)

专题命中 视频生成 :text-to-video(abstract);分类 cs.CV、cs.MM

Comments Accepted by TPAMI. The paper's homepage is https://github.com/HuiGuanLab/ms-sl-pp

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08568 2025-10-10 cs.RO cs.AI cs.CV 57%

NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos

Hongyu Li, Lingfeng Sun, Yafei Hu, Duy Ta, Jennifer Barry, George Konidaris, Jiahui Fu

机构 * Robotics and AI Institute(机器人与人工智能研究所) Brown University(布朗大学)

专题命中 视频生成 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频扩散模型 4 篇

2510.07345 2025-10-10 q-bio.QM cs.AI eess.IV 83%

Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model

Danush Kumar Venkatesh, Adam Schmidt, Muhammad Abdullah Jamal, Omid Mohareri

机构 * a Department of Translational Surgical Oncology, NCT/UCC Dresden, a partnership between DKFZ, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden,HZDR, Germany(a 转化外科肿瘤学部,NCT/UCC 德累斯顿,DKFZ、医学院和卡尔·古斯塔夫·卡尔斯医院、德累斯顿技术大学、HZDR 的联合体,德国) b Intuitive Surgical, Inc., Sunnyvale, CA, United States(b 直觉手术公司,美国加利福尼亚州 Sunnyvale)

专题命中 视频扩散模型 :video diffusion(title,abstract);video understanding(abstract);分类 eess.IV

Comments 29 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08530 2025-10-10 cs.GR cs.CV 70%

X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering

Zhitong Huang, Mohan Zhang, Renhan Wang, Rui Tang, Hao Zhu, Jing Liao

机构 * City University of Hong Kong(香港城市大学) WeChat, Tencent Inc(微信、腾讯公司) Manycore Tech Inc(很多核科技公司)

专题命中 视频扩散模型 :video generation(abstract);long video(abstract);分类 cs.CV

Comments Code, model, and dataset will be released at project page soon: https://luckyhzt.github.io/x2video

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07654 2025-10-10 cs.CV 57%

Once Is Enough: Lightweight DiT-Based Video Virtual Try-On via One-Time Garment Appearance Injection

Yanjie Pan, Qingdong He, Lidong Wang, Bo Peng, Mingmin Chi

机构 * School of computer science, Shanghai key laboratory of data science, Fudan University, China(计算机学院、数据科学重点实验室、复旦大学) Tencent Youtu Lab, China(腾讯优图实验室) Shanghai Ocean University, China(上海海洋大学)

专题命中 视频扩散模型 :video generation(abstract);分类 cs.CV

Comments 5 pages (including references), 4 figures. Code and models will be released upon publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15742 2025-10-10 cs.CV 57%

Uncertainty-Aware Diffusion Guided Refinement of 3D Scenes

Sarosij Bose, Arindam Dutta, Sayak Nag, Junge Zhang, Jiachen Li, Konstantinos Karydis, Amit K. Roy Chowdhury

机构 * University of California, Riverside, USA(加州大学河滨分校)

专题命中 视频扩散模型 :video diffusion(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频数据与评测 3 篇

2510.08559 2025-10-10 cs.CV cs.AI 79%

SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models

Andong Deng, Taojiannan Yang, Shoubin Yu, Lincoln Spencer, Mohit Bansal, Chen Chen, Serena Yeung-Levy, Xiaohan Wang

机构 * University of Central Florida(中央佛罗里达大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Stanford University(斯坦福大学)

专题命中 视频数据与评测 :video reasoning(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07550 2025-10-10 cs.CV cs.AI 79%

TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility

Saman Motamed, Minghao Chen, Luc Van Gool, Iro Laina

机构 * Visual Geometry Group, University of Oxford(视觉几何组,牛津大学)

专题命中 视频数据与评测 :video-language(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07441 2025-10-10 cs.CV 79%

DynamicEval: Rethinking Evaluation for Dynamic Text-to-Video Synthesis

Nithin C. Babu, Aniruddha Mahapatra, Harsh Rangwani, Rajiv Soundararajan, Kuldeep Kulkarni

机构 * Indian Institute of Science(印度科学研究院) Adobe Research(Adobe研究)

专题命中 视频数据与评测 :text-to-video(title,abstract);分类 cs.CV

Comments Preprint. Under review. 26 pages, 11 figures, 11 tables. Access the project page in https://nithincbabu7.github.io/DynamicEval

详情

展开后加载摘要…

URL PDF HTML 收藏