arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-11-20 至 2025-11-20 共收录 1 信号源:cs.CV, eess.IV, cs.MM

1. 动作与事件理解 1 篇

2404.03179 2025-11-20 cs.CV cs.MM cs.SD eess.AS 62%

UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization

Tiantian Geng, Teng Wang, Jinming Duan, Yanfu Zhang, Weili Guan, Feng Zheng, Ling shao

机构 * Department of Computer Science and Engineering, Southern University of Science and Technology(计算机科学与工程系,南方科技大学) School of Computer Science, University of Birmingham(计算机科学学院,伯明翰大学) Department of Computer Science, University of Hong Kong(计算机科学系,香港大学) Division of Informatics, Imaging and Data Sciences, University of Manchester(信息学、成像与数据科学系,曼彻斯特大学) William and Mary(威廉与玛丽学院) Harbin Institute of Technology(哈尔滨工业大学) UCAS-Terminus AI Lab, University of Chinese Academy of Sciences(中国科学院大学-Terminus AI实验室)

专题命中 动作与事件理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments Published on IEEE TPAMI

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 11, pp. 10280-10294, August 2025

详情

展开后加载摘要…

URL PDF HTML 收藏