arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 1388 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1388 篇

2504.02061 2025-04-04 cs.CV cs.MM cs.SD eess.AS 62%

Aligned Better, Listen Better for Audio-Visual Large Language Models

Yuxin Guo, Shuailei Ma, Shijie Ma, Xiaoyi Bao, Chen-Wei Xie, Kecheng Zheng, Tingyu Weng, Siyang Sun, Yun Zheng, Wei Zou

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19026 2025-02-27 eess.IV cs.AI cs.CV 62%

InternVQA: Advancing Compressed Video Quality Assessment with Distilling Large Foundation Model

Fengbin Guan, Zihao Yu, Yiting Lu, Xin Li, Zhibo Chen

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

Comments Accepted by ISCAS 2025(Lecture)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00365 2024-12-31 cs.AI cs.CV eess.IV 62%

Multimodal Fusion and Coherence Modeling for Video Topic Segmentation

Hai Yu, Chong Deng, Qinglin Zhang, Jiaqing Liu, Qian Chen, Wen Wang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05185 2024-12-12 cs.CV cs.LG cs.MM 62%

LinVT: Empower Your Image-level Large Language Model to Understand Videos

Lishuai Gao, Yujie Zhong, Yingsen Zeng, Haoxian Tan, Dengjie Li, Zheng Zhao

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06682 2024-10-14 cs.CV cs.CL eess.IV 62%

Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization

Changli Tang, Yixuan Li, Yudong Yang, Jimin Zhuang, Guangzhi Sun, Wei Li, Zujun Ma, Chao Zhang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21757 2024-09-13 cs.CV cs.MM 62%

Learning Video Context as Interleaved Multimodal Sequences

Kevin Qinghong Lin, Pengchuan Zhang, Difei Gao, Xide Xia, Joya Chen, Ziteng Gao, Jinheng Xie, Xuhong Xiao, Mike Zheng Shou

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments Accepted by ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03104 2024-08-13 cs.CV cs.CL cs.MM 62%

KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Hao Liang, Jiapeng Li, Tianyi Bai, Xijie Huang, Linzhuang Sun, Zhengren Wang, Conghui He, Bin Cui, Chong Chen, Wentao Zhang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02411 2024-07-04 cs.CV cs.CR cs.MM 62%

Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs

Jinmin Li, Kuofeng Gao, Yang Bai, Jingyun Zhang, Shu-Tao Xia

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments arXiv admin note: substantial text overlap with arXiv:2403.13507

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16301 2024-06-25 cs.CV cs.AI cs.MM 62%

UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos

Yuting Mei, Linli Yao, Qin Jin

专题命中 视频理解 :long video(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM International Conference on Multimedia Retrieval (ICMR'24)

Journal ref Proceedings of the 2024 International Conference on Multimedia Retrieval, May 2024, Pages 1034-1042

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04192 2024-04-02 cs.CV cs.AI cs.CL cs.MM 62%

Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models

Wei Han, Hui Chen, Min-Yen Kan, Soujanya Poria

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments 13 pages, 7 figures, accepted to Findings of NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.12968 2023-11-23 cs.CV cs.MM 62%

Improved Actor Relation Graph based Group Activity Recognition

Zijian Kuang, Xinran Tie

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Journal ref ICSM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00282 2023-11-13 cs.CV cs.IR cs.MM 62%

(Un)likelihood Training for Interpretable Embedding

Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan, Zhijian Hou

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments accepted in ACM Transactions on Information Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10227 2023-09-20 eess.IV cs.CV 62%

Learning Dynamic MRI Reconstruction with Convolutional Network Assisted Reconstruction Swin Transformer

Di Xu, Hengjie Liu, Dan Ruan, Ke Sheng

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

Comments MICCAI 2023 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13004 2023-08-28 cs.CV cs.AI cs.MM 62%

Spherical Vision Transformer for 360-degree Video Saliency Prediction

Mert Cokelek, Nevrez Imamoglu, Cagri Ozcinar, Erkut Erdem, Aykut Erdem

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments 12 pages, 4 figures, accepted to BMVC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09322 2023-08-21 cs.CV cs.AI cs.MM 62%

Audio-Visual Glance Network for Efficient Video Recognition

Muhammad Adi Nugroho, Sangmin Woo, Sumin Lee, Changick Kim

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03063 2023-08-08 cs.CV cs.MM 62%

M$^3$Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action Recognition

Hao Tang, Jun Liu, Shuanglin Yan, Rui Yan, Zechao Li, Jinhui Tang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13176 2023-06-26 cs.CV cs.LG eess.IV 62%

Key Frame Extraction with Attention Based Deep Neural Networks

Samed Arslan, Senem Tanberk

专题命中 视频理解 :long video(abstract);分类 cs.CV、eess.IV

Comments in Turkish language

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12037 2023-03-23 cs.CV cs.AI cs.LG cs.MM 62%

Causal Reasoning Meets Visual Representation Learning: A Prospective Study

Yang Liu, Yushen Wei, Hong Yan, Guanbin Li, Liang Lin

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments 35 pages, 14 figures. This work has been accepted by Machine Intelligence Research. The arxiv version is kept updating by adding more novel methods, datasets and insights. The official video interpretation of this paper can be referred at https://youtu.be/2lfNaTkcTHI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.04979 2023-03-17 cs.CV cs.LG cs.MM 62%

VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

Shen Yan, Tao Zhu, Zirui Wang, Yuan Cao, Mi Zhang, Soham Ghosh, Yonghui Wu, Jiahui Yu

专题命中 视频理解 :text-to-video(abstract);分类 cs.CV、cs.MM

Comments Tech report. arXiv v3: update text

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.04752 2022-11-15 eess.IV cs.CV 62%

Human Gaze Guided Attention for Surgical Activity Recognition

Abdishakour Awale, Duygu Sarikaya

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04212 2022-07-19 cs.CV cs.LG eess.IV 62%

AutoVideo: An Automated Video Action Recognition System

Daochen Zha, Zaid Pervaiz Bhat, Yi-Wei Chen, Yicheng Wang, Sirui Ding, Jiaben Chen, Kwei-Herng Lai, Mohammad Qazim Bhat, Anmoll Kumar Jain, Alfredo Costilla Reyes, Na Zou, Xia Hu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

Comments Accepted by IJCAI https://github.com/datamllab/autovideo

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.06931 2022-06-15 cs.CV cs.AI cs.MM 62%

Stand-Alone Inter-Frame Attention in Video Models

Fuchen Long, Zhaofan Qiu, Yingwei Pan, Ting Yao, Jiebo Luo, Tao Mei

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments CVPR 2022; Code is publicly available at: https://github.com/FuchenUSTC/SIFA

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03838 2022-03-09 cs.CV cs.AI cs.LG cs.MM 62%

Multi-Scale Self-Contrastive Learning with Hard Negative Mining for Weakly-Supervised Query-based Video Grounding

Shentong Mo, Daizong Liu, Wei Hu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.00629 2021-11-02 cs.MM cs.CV cs.IR cs.LG 62%

Distantly Supervised Semantic Text Detection and Recognition for Broadcast Sports Videos Understanding

Avijit Shah, Topojoy Biswas, Sathish Ramadoss, Deven Santosh Shah

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments 9 pages, 7 figures and 6 tables. To be published in the proceedings of ACM Multimedia 21, Industrial Track, held from October 20-24 in China

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.06637 2021-09-15 cs.CV cs.MM 62%

Multi-modal Representation Learning for Video Advertisement Content Structuring

Daya Guo, Zhaoyang Zeng

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02432 2021-08-06 cs.CV cs.MM 62%

Token Shift Transformer for Video Classification

Hao Zhang, Yanbin Hao, Chong-Wah Ngo

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments ACM Multimedia 2021, 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06567 2020-12-14 cs.CV cs.MM 62%

A Comprehensive Study of Deep Video Action Recognition

Yi Zhu, Xinyu Li, Chunhui Liu, Mohammadreza Zolfaghari, Yuanjun Xiong, Chongruo Wu, Zhi Zhang, Joseph Tighe, R. Manmatha, Mu Li

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments Technical report. Code and model zoo can be found at https://cv.gluon.ai/model_zoo/action_recognition.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.10260 2020-06-19 cs.CV cs.MM 62%

Video Moment Localization using Object Evidence and Reverse Captioning

Madhawa Vidanapathirana, Supriya Pandhre, Sonia Raychaudhuri, Anjali Khurana

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM

Comments 7 pages. 6 figures. For source code, refer https://github.com/madhawav/MML

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.07613 2020-01-22 cs.LG cs.CV eess.IV stat.ML 62%

Cut-Based Graph Learning Networks to Discover Compositional Structure of Sequential Video Data

Kyoung-Woon On, Eun-Sol Kim, Yu-Jung Heo, Byoung-Tak Zhang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

Comments 8 pages, 3 figures, Association for the Advancement of Artificial Intelligence (AAAI2020). arXiv admin note: substantial text overlap with arXiv:1907.01709

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.02793 2019-10-08 cs.CV cs.LG eess.IV 62%

ViP: Video Platform for PyTorch

Madan Ravi Ganesh, Eric Hofesmann, Nathan Louis, Jason Corso

专题命中 视频理解 :video understanding(abstract);分类 cs.CV、eess.IV

详情

展开后加载摘要…

URL PDF HTML 收藏