arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-10-08 至 2025-10-08 共收录 4 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 4 篇

2510.06077 2025-10-08 cs.CV cs.AI 83%

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

Mi Luo, Zihui Xue, Alex Dimakis, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) UC Berkeley(伯克利大学) Bespoke Labs(Bespoke实验室)

专题命中 视频理解 :video reasoning(title,abstract);video understanding(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2025, Project page: https://vision.cs.utexas.edu/projects/video-ver/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05836 2025-10-08 cs.CV 83%

Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow

Ruyang Liu, Shangkun Sun, Haoran Tang, Ge Li, Wei Gao

机构 * School of Electronic and Computer Engineering, Shenzhen Graduate School, 2 Peng Cheng LaboratoryPeking University(1 电子与计算机工程学院,深圳研究生院,2 深圳鹏城实验室,北京大学)

专题命中 视频理解 :video understanding(title,abstract);long video(abstract);分类 cs.CV

Comments Accepted to ICCV' 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03501 2025-10-08 cs.CV 83%

LV-MAE: Learning Long Video Representations through Masked-Embedding Autoencoders

Ilan Naiman, Emanuel Ben-Baruch, Oron Anschel, Alon Shoshan, Igor Kviatkovsky, Manoj Aggarwal, Gerard Medioni

机构 * Amazon(亚马逊)

专题命中 视频理解 :long video(title,abstract);video-language(abstract);分类 cs.CV

Comments Accepted to the International Conference on Computer Vision, ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15192 2025-10-08 cs.CV 57%

Leveraging Foundation Models for Multimodal Graph-Based Action Recognition

Fatemeh Ziaeetabar, Florentin Wörgötter

机构 * School of Mathematics, Statistics and Computer Science, College of Science, University of Tehran(数学、统计与计算机科学学院,科学学院,塔里斯坦大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏