arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6840 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1395 篇

2406.18070 2024-07-02 cs.CV 57%

EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Baoqi Pei, Guo Chen, Jilan Xu, Yuping He, Yicheng Liu, Kanghua Pan, Yifei Huang, Yali Wang, Tong Lu, Limin Wang, Yu Qiao

专题命中 视频理解 :video-language(abstract);分类 cs.CV

Comments Champion solutions in the EgoVis CVPR 2024 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15785 2024-06-28 cs.CV 57%

BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Ruyang Liu, Chen Li, Yixiao Ge, Ying Shan, Thomas H. Li, Ge Li

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13870 2024-06-27 cs.CV 57%

Splatter a Video: Video Gaussian Representation for Versatile Processing

Yang-Tian Sun, Yi-Hua Huang, Lin Ma, Xiaoyang Lyu, Yan-Pei Cao, Xiaojuan Qi

专题命中 视频理解 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.10636 2024-06-21 cs.CV cs.LG 57%

EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens

Sunil Hwang, Jaehong Yoon, Youngwan Lee, Sung Ju Hwang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted by ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13136 2024-06-21 cs.CV 57%

GVT2RPM: An Empirical Study for General Video Transformer Adaptation to Remote Physiological Measurement

Hao Wang, Euijoon Ahn, Jinman Kim

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Code is available at https://github.com/Dylan-H-Wang/facial-ai

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07191 2024-06-12 cs.CV 57%

MeMSVD: Long-Range Temporal Structure Capturing Using Incremental SVD

Ioanna Ntinou, Enrique Sanchez, Georgios Tzimiropoulos

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted to ICIP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16533 2024-06-12 cs.CV cs.AI cs.CL 57%

ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models

Avinash Madasu, Vasudev Lal

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04002 2024-06-10 cs.CV 57%

3rd Place Solution for PVUW Challenge 2024: Video Panoptic Segmentation

Ruipu Wu, Jifei Che, Han Li, Chengjing Wu, Ting Liu, Luoqi Liu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 3nd Place Solution for CVPR 2024 PVUW VPS Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00500 2024-06-04 cs.CV 57%

2nd Place Solution for PVUW Challenge 2024: Video Panoptic Segmentation

Biao Wu, Diankai Zhang, Si Gao, Chengjian Zheng, Shaoli Liu, Ning Wang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 2nd Place Solution for CVPR 2024 PVUW VPS Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17118 2024-05-27 cs.CV 57%

Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition

Jiaming Zhou, Hanjun Li, Kun-Yu Lin, Junwei Liang

专题命中 视频理解 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11487 2024-05-21 cs.CV 57%

"Previously on ..." From Recaps to Story Summarization

Aditya Kumar Singh, Dhruv Srivastava, Makarand Tapaswi

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments CVPR 2024; Project page: https://katha-ai.github.io/projects/recap-story-summ/

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05584 2024-05-10 cs.CV cs.AI 57%

A Survey on Backbones for Deep Video Action Recognition

Zixuan Tang, Youjun Zhao, Yuhang Wen, Mengyuan Liu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments This paper has been accepted by ICME workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09263 2024-05-09 cs.CV cs.AI 57%

Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection

Jin Yang, Ping Wei, Huan Li, Ziyang Ren

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04404 2024-05-08 cs.CV cs.AI cs.CL cs.LG 57%

Vision Mamba: A Comprehensive Survey and Taxonomy

Xiao Liu, Chenxu Zhang, Lei Zhang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments https://github.com/lx6c78/Vision-Mamba-A-Comprehensive-Survey-and-Taxonomy

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03272 2024-05-07 cs.CV 57%

WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Yuanhan Zhang, Kaichen Zhang, Bo Li, Fanyi Pu, Christopher Arif Setiadharma, Jingkang Yang, Ziwei Liu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15719 2024-04-26 cs.CV cs.AI 57%

HDBN: A Novel Hybrid Dual-branch Network for Robust Skeleton-based Action Recognition

Jinfu Liu, Baiqiao Yin, Jiaying Lin, Jiajun Wen, Yue Li, Mengyuan Liu

专题命中 视频理解 :video reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12060 2024-04-24 cs.CV cs.CL 57%

VideoXum: Cross-modal Visual and Textural Summarization of Videos

Jingyang Lin, Hang Hua, Ming Chen, Yikang Li, Jenhao Hsiao, Chiuman Ho, Jiebo Luo

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments 13 pages, 7 figures

Journal ref IEEE Transactions on Multimedia, VOL. 26 (2024) 5548-5560

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.11470 2024-04-18 cs.CV 57%

Exploring Missing Modality in Multimodal Egocentric Datasets

Merey Ramazanova, Alejandro Pardo, Humam Alwassel, Bernard Ghanem

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09220 2024-04-16 cs.CV 57%

A Survey on Open-Vocabulary Detection and Segmentation: Past, Present, and Future

Chaoyang Zhu, Long Chen

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08229 2024-04-15 cs.CV 57%

Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis

Maged Shoman, Dongdong Wang, Armstrong Aboah, Mohamed Abdel-Aty

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03398 2024-04-05 cs.CV 57%

Scaling Up Video Summarization Pretraining with Large Language Models

Dawit Mureja Argaw, Seunghyun Yoon, Fabian Caba Heilbron, Hanieh Deilamsalehy, Trung Bui, Zhaowen Wang, Franck Dernoncourt, Joon Son Chung

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments Accepted to CVPR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01297 2024-04-02 cs.CV 57%

Streaming Dense Video Captioning

Xingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan, Austin Myers, Xuehan Xiong, Arsha Nagrani, Cordelia Schmid

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments CVPR 2024. Code is available at https://github.com/google-research/scenic/tree/main/scenic/projects/streaming_dvc

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00901 2024-04-02 cs.CV 57%

Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding

Syed Talal Wasim, Muzammal Naseer, Salman Khan, Ming-Hsuan Yang, Fahad Shahbaz Khan

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14326 2024-03-28 cs.MM 57%

Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic Segmentation

Mingxuan Yan, Yi Wang, Xuedou Xiao, Zhiqing Luo, Jianhua He, Wei Wang

专题命中 视频理解 :video understanding(abstract);分类 cs.MM

Comments Accepted by ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09506 2024-03-26 cs.CV cs.AI cs.LG 57%

Don't Judge by the Look: Towards Motion Coherent Video Representation

Yitian Zhang, Yue Bai, Huan Wang, Yizhou Wang, Yun Fu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted by ICLR2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04940 2024-03-11 cs.CV cs.AI cs.LG q-bio.NC 57%

A spatiotemporal style transfer algorithm for dynamic visual stimulus generation

Antonino Greco, Markus Siegel

专题命中 视频理解 :video generation(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.19479 2024-03-01 cs.CV 57%

Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, Sergey Tulyakov

专题命中 视频理解 :video generation(abstract);分类 cs.CV

Comments CVPR 2024. Project Page: https://snap-research.github.io/Panda-70M

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02574 2024-02-07 cs.CV cs.LG 57%

Spatio-temporal Prompting Network for Robust Video Feature Extraction

Guanxiong Sun, Chi Wang, Zhaoyu Zhang, Jiankang Deng, Stefanos Zafeiriou, Yang Hua

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Journal ref 2023 International Conference on Computer Vision (ICCV) 13541-13551

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10254 2024-01-22 cs.CV cs.LG 57%

Beyond the Frame: Single and mutilple video summarization method with user-defined length

Vahid Ahmadi Kalkhorani, Qingquan Zhang, Guanqun Song, Ting Zhu

专题命中 视频理解 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.10702 2024-01-22 cs.CV 57%

ClawCraneNet: Leveraging Object-level Relation for Text-based Video Segmentation

Chen Liang, Yu Wu, Yawei Luo, Yi Yang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Extended version published in https://ieeexplore.ieee.org/abstract/document/10083244

详情

展开后加载摘要…

URL PDF HTML 收藏