arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-08-18 至 2025-08-18 共收录 1 信号源:cs.CV, eess.IV, cs.MM

1. 视频数据与评测 1 篇

2508.10922 2025-08-18 cs.CV 70%

A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu, Wei Liu, Ye Liu, Meng Liu, Liqiang Nie, Zhouchen Lin, Chang Wen Chen

机构 * School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) School of Computer Science and Technology, Shandong Jianzhu University(山东建筑大学计算机科学与技术学院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 视频数据与评测 :video understanding(abstract);video-language(abstract);分类 cs.CV

Comments 20 pages,6 figures,survey

详情

展开后加载摘要…

URL PDF HTML 收藏