arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-17 至 2025-10-17 共收录 2 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 2 篇

2510.14032 2025-10-17 cs.CV 89%

Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding

Xiaoqian Shen, Wenxuan Zhang, Jun Chen, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology(卡布斯大学)

专题命中 视频理解 :video understanding(title,abstract);long video(title,abstract);video language model(abstract);分类 cs.CV

Comments NeurIPS 2025 (Spotlight). Webpage at https://xiaoqian-shen.github.io/Vgent

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13891 2025-10-17 cs.LG cs.AI 86%

K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding

Yifeng Yao, Yike Yun, Jing Wang, Huishuai Zhang, Dongyan Zhao, Ke Tian, Zhihao Wang, Minghui Qiu, Tao Wang

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) Bytedance(字节跳动)

专题命中 视频理解 :video understanding(title,abstract);long video(title)

详情

展开后加载摘要…

URL PDF HTML 收藏