arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 1388 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1388 篇

2307.16715 2023-08-21 cs.CV 74%

UniVTG: Towards Unified Video-Language Temporal Grounding

Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick, Difei Gao, Alex Jinpeng Wang, Rui Yan, Mike Zheng Shou

专题命中 视频理解 :video-language(title);分类 cs.CV

Comments Accepted by ICCV 2023. 16 pages, 10 figures, 13 tables. Code: https://github.com/showlab/UniVTG

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.09431 2023-06-19 cs.MM 74%

Towards Long Form Audio-visual Video Understanding

Wenxuan Hou, Guangyao Li, Yapeng Tian, Di Hu

专题命中 视频理解 :video understanding(title);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09539 2023-05-17 cs.CV 74%

Learning Higher-order Object Interactions for Keypoint-based Video Understanding

Yi Huang, Asim Kadav, Farley Lai, Deep Patel, Hans Peter Graf

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments SRVU - ICCV' 2021 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13402 2023-04-27 cs.CV 74%

Concept Graph Neural Networks for Surgical Video Understanding

Yutong Ban, Jennifer A. Eckhoff, Thomas M. Ward, Daniel A. Hashimoto, Ozanan R. Meireles, Daniela Rus, Guy Rosman

专题命中 视频理解 :video understanding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00325 2023-04-25 cs.CV 74%

SVT: Supertoken Video Transformer for Efficient Video Understanding

Chenbin Pan, Rui Hou, Hanchao Yu, Qifan Wang, Senem Velipasalar, Madian Khabsa

专题命中 视频理解 :video understanding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.18230 2023-04-03 cs.CV 74%

Procedure-Aware Pretraining for Instructional Video Understanding

Honglu Zhou, Roberto Martín-Martín, Mubbasir Kapadia, Silvio Savarese, Juan Carlos Niebles

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.11329 2022-07-26 cs.CV 74%

Video Swin Transformers for Egocentric Video Understanding @ Ego4D Challenges 2022

Maria Escobar, Laura Daza, Cristina González, Jordi Pont-Tuset, Pablo Arbeláez

专题命中 视频理解 :video understanding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01975 2022-07-21 cs.CV 74%

Federated Self-supervised Learning for Video Understanding

Yasar Abbas Ur Rehman, Yan Gao, Jiajun Shen, Pedro Porto Buarque de Gusmao, Nicholas Lane

专题命中 视频理解 :video understanding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.04710 2021-12-10 cs.CV 74%

Auto-X3D: Ultra-Efficient Video Understanding via Finer-Grained Neural Architecture Search

Yifan Jiang, Xinyu Gong, Junru Wu, Humphrey Shi, Zhicheng Yan, Zhangyang Wang

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments Accepted by WACV'2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.05095 2021-06-10 cs.CV 74%

Is Space-Time Attention All You Need for Video Understanding?

Gedas Bertasius, Heng Wang, Lorenzo Torresani

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments Accepted to ICML 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.09496 2021-04-20 cs.CV 74%

Temporal Query Networks for Fine-grained Video Understanding

Chuhan Zhang, Ankush Gupta, Andrew Zisserman

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments Accepted to CVPR 2021(Oral). Project page: http://www.robots.ox.ac.uk/~vgg/research/tqn/

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.00830 2020-08-03 cs.CV 74%

Temporal Aggregate Representations for Long-Range Video Understanding

Fadime Sener, Dipika Singhania, Angela Yao

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments ECCV 2020, European Conference on Computer Vision

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.03857 2019-06-13 cs.CV 74%

UniDual: A Unified Model for Image and Video Understanding

Yufei Wang, Du Tran, Lorenzo Torresani

专题命中 视频理解 :video understanding(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.05038 2019-04-19 cs.CV 74%

Long-Term Feature Banks for Detailed Video Understanding

Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krähenbühl, Ross Girshick

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments Code and models are available at https://github.com/facebookresearch/video-long-term-feature-banks

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.09834 2018-11-27 cs.CV cs.LG stat.ML 74%

Efficient Video Understanding via Layered Multi Frame-Rate Analysis

Ziyao Tang, Yongxi Lu, Tara Javidi

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.00413 2018-04-03 cs.CV 74%

End-to-End Learning of Motion Representation for Video Understanding

Lijie Fan, Wenbing Huang, Chuang Gan, Stefano Ermon, Boqing Gong, Junzhou Huang

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments CVPR 2018 spotlight. The first two authors contributed equally to this paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.04488 2017-06-15 cs.CV 74%

Large-Scale YouTube-8M Video Understanding with Deep Neural Networks

Manuk Akopyan, Eshsou Khashba

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments 6 pages, 5 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1206.5065 2013-03-04 cs.CV 74%

A generic framework for video understanding applied to group behavior recognition

Sofia Zaidenberg, Bernard Boulay, François Bremond

专题命中 视频理解 :video understanding(title);分类 cs.CV

Comments (20/03/2012)

Journal ref 9th IEEE International Conference on Advanced Video and Signal-Based Surveillance (AVSS 2012) (2012) 136 -142

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22967 2025-10-21 cs.CV cs.LG cs.MM 73%

ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment

Amir Aghdam, Vincent Tao Hu, Björn Ommer

机构 * Department of Computer Science, Temple University(Temple大学计算机科学系) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 视频理解 :video understanding(abstract);video-language(abstract);分类 cs.CV、cs.MM

Comments Accepted to TMLR 2025 - Project page: https://amir-aghdam.github.io/act-align/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13809 2024-06-21 cs.MM cs.CV cs.IR 73%

Towards Holistic Language-video Representation: the language model-enhanced MSR-Video to Text Dataset

Yuchen Yang, Yingxuan Duan

专题命中 视频理解 :video understanding(abstract);video-language(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00220 2023-12-04 cs.MM cs.CL cs.CV 73%

Multi-Modal Video Topic Segmentation with Dual-Contrastive Domain Adaptation

Linzi Xing, Quan Tran, Fabian Caba, Franck Dernoncourt, Seunghyun Yoon, Zhaowen Wang, Trung Bui, Giuseppe Carenini

专题命中 视频理解 :video understanding(abstract);long video(abstract);分类 cs.CV、cs.MM

Comments Accepted at the 30th International Conference on Multimedia Modeling (MMM 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.14895 2022-05-31 cs.CV cs.CL cs.MM 73%

From Representation to Reasoning: Towards both Evidence and Commonsense Reasoning for Video Question-Answering

Jiangtong Li, Li Niu, Liqing Zhang

专题命中 视频理解 :video understanding(abstract);video reasoning(abstract);分类 cs.CV、cs.MM

Comments To appear in CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09758 2023-10-27 cs.CV cs.CL 72%

A Video Is Worth 4096 Tokens: Verbalize Videos To Understand Them In Zero Shot

Aanisha Bhattacharya, Yaman K Singla, Balaji Krishnamurthy, Rajiv Ratn Shah, Changyou Chen

专题命中 视频理解 :video understanding(abstract,comments);long video(abstract);分类 cs.CV

Comments Accepted to EMNLP-23 TL;DR: Video understanding lags far behind NLP; LLMs excel in zero-shot. Our approach utilizes LLMs to verbalize videos, creating stories for zero-shot video understanding. This yields state-of-the-art results across five datasets, covering fifteen tasks

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25584 2026-04-29 cs.AI 71%

DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding

DualFact+: 一种用于过程视频理解的多模态事实验证框架

Cennet Oguz, Yasser Hamidullah, Josef van Genabith, Simon Ostermann

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) Saarland Informatics Campus(萨尔兰信息学校区)

专题命中 视频理解 :video understanding(title)

AI总结 DualFact+通过双层多模事实验证框架,针对过程视频描述中的概念事实和上下文事实进行评估,揭示了多模态事实 grounding 的挑战。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04829 2022-10-11 cs.CL 71%

Hierarchical3D Adapters for Long Video-to-text Summarization

Pinelopi Papalampidi, Mirella Lapata

专题命中 视频理解 :long video(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08167 2026-08-11 cs.CV cs.LG 新提交 70%

Wiener Representation Filtering for VLM Hallucination Suppression

用于抑制视觉语言模型幻觉的维纳表示滤波

Ameen Ali, Tamim Zoabi, Lidor Brami, Lior Wolf

机构 * Tel Aviv University(特拉维夫大学)

专题命中 视频理解 :video understanding(abstract);video reasoning(abstract);分类 cs.CV

AI总结 该研究提出一种无需训练的事后维纳表示滤波技术,通过离线校准校正视觉语言模型深层前馈输出投影,可在保持运行速度的同时降低其对象幻觉,在多类模型及基准上均验证了有效性与通用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03083 2026-08-05 cs.CV cs.CL 新提交 70%

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models

GSTEP:面向高效视频大语言模型的全局时空密度驱动视觉令牌剪枝

Mengjie Zhang, Qihui Zhu, Tao Zhang, Shuangwu Chen, Huihuang Qin, Yu Guo, Shenghao Ye, Zijian Wen, Yunpeng Hou, Dong Jin, Xiaobin Tan, Huasen He, Jian Yang

专题命中 视频理解 :video understanding(abstract);long video(abstract);分类 cs.CV

AI总结 本文提出即插即用的GSTEP剪枝框架,解决现有视频令牌剪枝方法的局部剪枝缺陷,在多个VideoLLMs上实现良好的准确率-效率权衡,在LLaVA-OneVision-7B上剪去75%视觉令牌,保留100.2%原始性能并获1.17倍端到端加速。

Comments 4 figures, accepted to ACM MM 26'

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23265 2026-08-04 cs.CV 版本更新 70%

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation

WaveZip:用于视频令牌压缩的小波驱动时空解耦

Yuhui Zeng, Wang Chen, Jinfa Huang, Tianyu Xie, Yongdong Luo, Jiayi Ji, Xiawu Zheng, jiebo Luo

机构 * Media Analytics and Computing Lab, Xiamen University(厦门大学媒体分析与计算实验室) Department of Computer Science, University of Rochester(罗切斯特大学计算机科学系)

专题命中 视频理解 :video understanding(abstract);long video(abstract);分类 cs.CV

AI总结 针对长视频理解中视觉令牌计算成本高的问题,提出WaveZip框架,利用离散小波变换分离时空信号,无需特定任务训练,能集成到LVLMs提升推理效率,在长视频理解基准测试中表现优异。

Comments 13 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.17994 2026-07-21 cs.CV cs.AI 新提交 70%

HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization

HAS:用于多模态大语言模型视频摘要的高亮引导注意力转向

Rui Chu, Yingjie Lao

机构 * Tufts University(塔夫茨大学)

专题命中 视频理解 :video generation(abstract);video understanding(abstract);分类 cs.CV

AI总结 针对视频摘要,提出HAS方法,通过全局找连续帧级高亮分布并作为注意力转向向量,让多模态大语言模型在推理时更关注高亮帧,避免丢失信息,在多种基准测试中性能出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10607 2026-07-16 cs.CV 版本更新 70%

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation

跟踪并描述任何运动:通过轨迹条件生成实现开放词汇时空字幕

Bishoy Galoaa, Sarah Ostadabbas

机构 * Northeastern University(东北大学)

专题命中 视频理解 :video understanding(abstract);video-language(abstract);分类 cs.CV

AI总结 研究提出TCAM框架,无需文本查询和区域提示,通过字幕感知重采样器在点粒度结合跟踪与语言,用现有分割注释训练,能描述视频中运动、定位时间及轨迹,优于密集视频字幕基线,匹配相关方法,为运动驱动视频理解提供新途径。

详情

展开后加载摘要…

URL PDF HTML 收藏