arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-09-30 至 2025-09-30 共收录 5 信号源:cs.CV, eess.IV, cs.MM

1. 长视频与时序推理 5 篇

2509.25161 2025-09-30 cs.CV 88%

Rolling Forcing: Autoregressive Long Video Diffusion in Real Time

Kunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan, Shijian Lu

机构 * Nanyang Technological University(南洋理工大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室)

专题命中 长视频与时序推理 :long video(title,abstract);video diffusion(title);video generation(abstract);分类 cs.CV

Comments Project page: https://kunhao-liu.github.io/Rolling_Forcing_Webpage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02850 2025-09-30 cs.CV 86%

METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding

Mengyue Wang, Shuo Chen, Kristian Kersting, Volker Tresp, Yunpu Ma

专题命中 长视频与时序推理 :long video(title,abstract);video understanding(title);分类 cs.CV

Comments EMNLP 2025; 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24943 2025-09-30 cs.CV 83%

Perceive, Reflect and Understand Long Video: Progressive Multi-Granular Clue Exploration with Interactive Agents

Jiahua Li, Kun Wei, Zhe Xu, Zibo Su, Xu Yang, Cheng Deng

机构 * School of Electronic Engineering, Xidian University(西安电子科技大学电子工程学院) Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港理工大学计算机科学与工程系) College of Computer and Information, Hohai University(河海大学计算机与信息学院)

专题命中 长视频与时序推理 :long video(title,abstract);video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24871 2025-09-30 cs.CV 79%

StreamForest: Efficient Online Video Understanding with Persistent Event Memory

Xiangyu Zeng, Kefan Qiu, Qingyu Zhang, Xinhao Li, Jing Wang, Jiaxin Li, Ziang Yan, Kun Tian, Meng Tian, Xinhai Zhao, Yi Wang, Limin Wang

机构 * Nanjing University(南京大学) Shanghai AI Laboratory(上海人工智能实验室) Zhejiang University(浙江大学) Noah’s Ark Lab, Huawei(华为诺亚实验室) Yinwang Intelligent Tech.(云网智能科技)

专题命中 长视频与时序推理 :video understanding(title,abstract);分类 cs.CV

Comments Accepted as a Spotlight at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25183 2025-09-30 cs.CV 57%

PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos

Ting-Hsuan Liao, Haowen Liu, Yiran Xu, Songwei Ge, Gengshan Yang, Jia-Bin Huang

机构 * University of Maryland College Park(马里兰大学学院公园分校)

专题命中 长视频与时序推理 :long video(abstract);分类 cs.CV

Comments SIGGRAPH Asia 2025. Project page:https://pad3r.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏