arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 1388 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 1388 篇

2111.13196 2022-06-22 cs.CV 57%

SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning

Kevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed, Zhe Gan, Zicheng Liu, Yumao Lu, Lijuan Wang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.10990 2022-06-20 cs.CV 57%

Learning To Recognize Procedural Activities with Distant Supervision

Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, Lorenzo Torresani

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments CVPR 2022. Code will be released here https://github.com/facebookresearch/video-distant-supervision

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06915 2022-06-13 cs.CV 57%

Object-Region Video Transformers

Roei Herzig, Elad Ben-Avraham, Karttikeya Mangalam, Amir Bar, Gal Chechik, Anna Rohrbach, Trevor Darrell, Amir Globerson

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02002 2022-06-07 cs.CV cs.LG 57%

CVNets: High Performance Library for Computer Vision

Sachin Mehta, Farzad Abdolhosseini, Mohammad Rastegari

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.04288 2022-06-01 cs.CV cs.LG 57%

Multiview Transformers for Video Recognition

Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, Cordelia Schmid

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments CVPR 2022; arXiv v4: update results on Epic-Kitchens-100

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01818 2022-05-06 cs.LG cs.AI cs.CL cs.CV eess.AS 57%

i-Code: An Integrative and Composable Multimodal Learning Framework

Ziyi Yang, Yuwei Fang, Chenguang Zhu, Reid Pryzant, Dongdong Chen, Yu Shi, Yichong Xu, Yao Qian, Mei Gao, Yi-Ling Chen, Liyang Lu, Yujia Xie, Robert Gmyr, Noel Codella, Naoyuki Kanda, Bin Xiao, Lu Yuan, Takuya Yoshioka, Michael Zeng, Xuedong Huang

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12408 2022-04-27 cs.CV 57%

MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval

Yuying Ge, Yixiao Ge, Xihui Liu, Alex Jinpeng Wang, Jianping Wu, Ying Shan, Xiaohu Qie, Ping Luo

专题命中 视频理解 :text-to-video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12293 2022-04-27 cs.CV 57%

Contrastive Language-Action Pre-training for Temporal Localization

Mengmeng Xu, Erhan Gundogdu, Maksim Lapin, Bernard Ghanem, Michael Donoser, Loris Bazzani

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 18 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.14104 2022-04-27 cs.CV 57%

Can An Image Classifier Suffice For Action Recognition?

Quanfu Fan, Chun-Fu, Chen, Rameswar Panda

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10916 2022-04-26 cs.CV cs.LG 57%

Revealing Occlusions with 4D Neural Fields

Basile Van Hoorick, Purva Tendulkar, Didac Suris, Dennis Park, Simon Stent, Carl Vondrick

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments CVPR 2022 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08874 2022-04-20 cs.CV 57%

Less than Few: Self-Shot Video Instance Segmentation

Pengwan Yang, Yuki M. Asano, Pascal Mettes, Cees G. M. Snoek

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments 25 pages, 5 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08671 2022-04-20 cs.CV 57%

ActAR: Actor-Driven Pose Embeddings for Video Action Recognition

Soufiane Lamghari, Guillaume-Alexandre Bilodeau, Nicolas Saunier

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03141 2022-04-08 cs.CV cs.AI 57%

Adversarial Machine Learning Attacks Against Video Anomaly Detection Systems

Furkan Mumcu, Keval Doshi, Yasin Yilmaz

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.11297 2022-04-05 cs.CV cs.LG 57%

TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Michael S. Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, Anelia Angelova

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments This is the full version of the paper, extending its conference paper at NeurIPS 2021. Version 1.1 of the code is released

Journal ref NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.09212 2022-04-01 cs.CV cs.AI 57%

Long-Short Temporal Contrastive Learning of Video Transformers

Jue Wang, Gedas Bertasius, Du Tran, Lorenzo Torresani

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted in CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.07974 2022-03-31 cs.CV 57%

TCLR: Temporal Contrastive Learning for Video Representation

Ishan Dave, Rohit Gupta, Mamshad Nayeem Rizve, Mubarak Shah

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted to Computer Vision and Image Understanding (CVIU) Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15205 2022-03-30 cs.CV cs.CR cs.LG 57%

SPAct: Self-supervised Privacy Preservation for Action Recognition

Ishan Rajendrakumar Dave, Chen Chen, Mubarak Shah

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments CVPR-2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15086 2022-03-30 cs.CV 57%

X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval

Satya Krishna Gorti, Noel Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, Guangwei Yu

专题命中 视频理解 :video reasoning(abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.04850 2022-03-18 cs.CV 57%

Bridging Video-text Retrieval with Multiple Choice Questions

Yuying Ge, Yixiao Ge, Xihui Liu, Dian Li, Ying Shan, Xiaohu Qie, Ping Luo

专题命中 视频理解 :text-to-video(abstract);分类 cs.CV

Comments Accepted by CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.02573 2022-03-08 cs.CV cs.AI cs.LG 57%

Show Me What and Tell Me How: Video Synthesis via Multimodal Conditioning

Ligong Han, Jian Ren, Hsin-Ying Lee, Francesco Barbieri, Kyle Olszewski, Shervin Minaee, Dimitris Metaxas, Sergey Tulyakov

专题命中 视频理解 :video generation(abstract);分类 cs.CV

Comments Accepted to CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.10828 2022-02-23 cs.CV 57%

Exploiting long-term temporal dynamics for video captioning

Yuyu Guo, Jingqiu Zhang, Lianli Gao

专题命中 视频理解 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.09979 2022-02-22 cs.CL cs.CV 57%

Audio Visual Scene-Aware Dialog Generation with Transformer-based Video Representations

Yoshihiro Yamazaki, Shota Orihashi, Ryo Masumura, Mihiro Uchida, Akihiko Takashima

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted at DSTC10 Workshop at AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.12086 2022-02-16 cs.CV 57%

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Junnan Li, Dongxu Li, Caiming Xiong, Steven Hoi

专题命中 视频理解 :video-language(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.09153 2022-01-25 cs.CV cs.AI 57%

An Integrated Approach for Video Captioning and Applications

Soheyla Amirian, Thiab R. Taha, Khaled Rasheed, Hamid R. Arabnia

专题命中 视频理解 :long video(abstract);分类 cs.CV

Comments The 2021 World Congress in Computer Science, Computer Engineering, and Applied Computing (CSCE'21), IEEE, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02359 2022-01-11 cs.CL cs.CV 57%

O2NA: An Object-Oriented Non-Autoregressive Approach for Controllable Video Captioning

Fenglin Liu, Xuancheng Ren, Xian Wu, Bang Yang, Shen Ge, Yuexian Zou, Xu Sun

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted by Findings of ACL 2021 (The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing)

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08913 2021-12-21 cs.CV 57%

Contrastive Spatio-Temporal Pretext Learning for Self-supervised Video Representation

Yujia Zhang, Lai-Man Po, Xuyuan Xu, Mengyang Liu, Yexin Wang, Weifeng Ou, Yuzhi Zhao, Wing-Yin Yu

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Accepted by AAAI 2022, Preprint version with Appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.06423 2021-12-17 cs.CV 57%

Zero-Shot Action Recognition in Videos: A Survey

Valter Estevam, Helio Pedrini, David Menotti

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.03803 2021-12-09 cs.CV 57%

Suppressing Static Visual Cues via Normalizing Flows for Self-Supervised Video Representation Learning

Manlin Zhang, Jinpeng Wang, Andy J. Ma

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments AAAI2022. v2: Add supplementary

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.01038 2021-12-03 cs.CV 57%

Stacked Temporal Attention: Improving First-person Action Recognition by Emphasizing Discriminative Clips

Lijin Yang, Yifei Huang, Yusuke Sugano, Yoichi Sato

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

Comments BMVC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.14785 2021-12-02 cs.CV cs.CL 57%

A Comprehensive Review of the Video-to-Text Problem

Jesus Perez-Martin, Benjamin Bustos, Silvio Jamil F. Guimarães, Ivan Sipiran, Jorge Pérez, Grethel Coello Said

专题命中 视频理解 :text-to-video(abstract);分类 cs.CV

Comments 66 pages, 6 figures. Accepted by Artificial Intelligence Review

详情

展开后加载摘要…

URL PDF HTML 收藏