arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2025-12-04 至 2025-12-04 共收录 2 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 2 篇

2512.03500 2025-12-04 cs.CV 86%

EEA: Exploration-Exploitation Agent for Long Video Understanding

EEA:长视频理解的探索-利用代理

Te Yang, Xiangyu Zhu, Bo Wang, Quan Chen, Peng Jiang, Zhen Lei

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Kuaishou Technology(快手科技) Centre for Artificial Intelligence and Robotics, HKISI, Chinese Academy of Sciences(人工智能与机器人中心,HKISI,中国科学院)

专题命中 视频理解 :video understanding(title,abstract);long video(title);分类 cs.CV

AI总结 EEA通过语义引导的分层树搜索实现长视频理解中的探索与利用平衡,结合视觉语言模型和语义先验提升视频分析的效率与精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04025 2025-12-04 cs.CV cs.AI cs.LG 79%

PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation

PSA: 基于金字塔稀疏注意力的高效视频理解和生成

Xiaolong Li, Youping Gu, Xi Lin, Weijie Wang, Bohan Zhuang

机构 * ZIP Lab, Zhejiang University(浙江大学浙大信息实验室)

专题命中 视频理解 :video understanding(title,abstract);分类 cs.CV

AI总结 PSA通过多级池化键值表示实现高效视频理解和生成,相比现有稀疏注意力方法在效率和质量上表现更优。

Comments Tech report

详情

展开后加载摘要…

URL PDF HTML 收藏