arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-27 至 2025-10-27 共收录 9 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 2 篇

2506.03525 2025-10-27 cs.CV cs.AI cs.CL 83%

Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning

Daeun Lee, Jaehong Yoon, Jaemin Cho, Mohit Bansal

机构 * UNC Chapel Hill(UNC夏洛特山分校) Nanyang Technological University(南洋理工大学)

专题命中 视频理解 :video reasoning(title,abstract);video understanding(abstract);分类 cs.CV

Comments Project website: https://video-skill-cot.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21026 2025-10-27 cs.RO 50%

HRT1: One-Shot Human-to-Robot Trajectory Transfer for Mobile Manipulation

Sai Haneesh Allu, Jishnu Jaykumar P, Ninad Khargonkar, Tyler Summers, Jian Yao, Yu Xiang

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校) XPeng(小鹏)

专题命中 视频理解 :video understanding(abstract)

Comments 14 pages, 11 figures and 3 tables. Project page is available at \url{https://irvlutd.github.io/HRT1/}

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频生成 4 篇

2510.21696 2025-10-27 cs.CV 83%

BachVid: Training-Free Video Generation with Consistent Background and Character

Han Yan, Xibin Song, Yifu Wang, Hongdong Li, Pan Ji, Chao Ma

机构 * MoE Key Lab of Artificial, AI Institute, Shanghai Jiao Tong University(人工智能研究院,上海交通大学) Vertex Lab(Vertex实验室) Australian National University(澳大利亚国立大学)

专题命中 视频生成 :video generation(title,abstract);text-to-video(abstract);分类 cs.CV

Comments Project page: https://wolfball.github.io/bachvid

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16500 2025-10-27 cs.CV 83%

RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation

Tianyi Yan, Wencheng Han, Xia Zhou, Xueyang Zhang, Kun Zhan, Cheng-zhong Xu, Jianbing Shen

机构 * SKL-IOTSC, Computer and Information Science, University of Macau(澳门大学计算机与信息科学学院) Li Auto Inc(利汽车公司)

专题命中 视频生成 :video generation(title,abstract);video diffusion(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20888 2025-10-27 cs.CV cs.AI 83%

Video-As-Prompt: Unified Semantic Control for Video Generation

Yuxuan Bian, Xin Chen, Zenan Li, Tiancheng Zhi, Shen Sang, Linjie Luo, Qiang Xu

机构 * Intelligent Creation Lab, ByteDance(字节跳动智能创作实验室) The Chinese University of Hong Kong(香港中文大学)

专题命中 视频生成 :video generation(title,abstract);video diffusion(abstract);分类 cs.CV

Comments Website: https://bytedance.github.io/Video-As-Prompt

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21491 2025-10-27 cs.CV 83%

Frame In-N-Out: Unbounded Controllable Image-to-Video Generation

Boyang Wang, Xuweiyi Chen, Matheus Gadelha, Zezhou Cheng

机构 * University of Virginia(弗吉尼亚大学) Adobe Research(Adobe研究)

专题命中 视频生成 :video generation(title,abstract);video diffusion(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 长视频与时序推理 2 篇

2506.15745 2025-10-27 eess.IV cs.LG 83%

InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding

Minsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung Chang

专题命中 长视频与时序推理 :video understanding(title,abstract);long video(abstract);分类 eess.IV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06485 2025-10-27 cs.CV cs.AI cs.CL 79%

Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning

Ziyang Wang, Jaehong Yoon, Shoubin Yu, Md Mohaiminul Islam, Gedas Bertasius, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Nanyang Technological University(南洋理工大学)

专题命中 长视频与时序推理 :video reasoning(title,abstract);分类 cs.CV

Comments EMNLP 2025. The first two authors contributed equally. Project page: https://sites.google.com/cs.unc.edu/videorts2025/

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 视频数据与评测 1 篇

2510.21406 2025-10-27 cs.CV 57%

MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence

Yue Feng, Jinwei Hu, Qijia Lu, Jiawei Niu, Li Tan, Shuo Yuan, Ziyi Yan, Yizhen Jia, Qingzhi He, Shiping Ge, Ethan Q. Chen, Wentong Li, Limin Wang, Jie Qin

机构 * MoE Key Laboratory of Brain-Machine Intelligence Technology, College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(脑机智能技术MoE实验室,人工智能学院,南京航空航天大学) Nanjing University(南京大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2025 D&B Track

详情

展开后加载摘要…

URL PDF HTML 收藏