arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

共收录 6790 信号源:cs.CV, eess.IV, cs.MM

1. 视频数据与评测 578 篇

2312.15670 2023-12-27 cs.CV 57%

Open-Vocabulary Video Relation Extraction

Wentao Tian, Zheng Wang, Yuqian Fu, Jingjing Chen, Lechao Cheng

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments accpeted by AAAI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00213 2023-12-04 cs.CV cs.AI 57%

Consistent Video-to-Video Transfer Using Synthetic Dataset

Jiaxin Cheng, Tianjun Xiao, Tong He

专题命中 视频数据与评测 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17982 2023-12-01 cs.CV 57%

VBench: Comprehensive Benchmark Suite for Video Generative Models

Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, Ziwei Liu

专题命中 视频数据与评测 :video generation(abstract);分类 cs.CV

Comments Equal contributions from first four authors. Project page: https://vchitect.github.io/VBench-project/ Code: https://github.com/Vchitect/VBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05526 2023-11-28 cs.CV 57%

Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignment

Zihui Xue, Kristen Grauman

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments Accepted by NeurIPS 2023, Project website: https://vision.cs.utexas.edu/projects/AlignEgoExo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13903 2023-11-10 cs.CL cs.CV 57%

Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought

Vaishnavi Himakunthala, Andy Ouyang, Daniel Rose, Ryan He, Alex Mei, Yujie Lu, Chinmay Sonar, Michael Saxon, William Yang Wang

专题命中 视频数据与评测 :video reasoning(abstract);分类 cs.CV

Comments Accepted to the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00556 2023-11-02 cs.CV 57%

ProBio: A Protocol-guided Multimodal Dataset for Molecular Biology Lab

Jieming Cui, Ziren Gong, Baoxiong Jia, Siyuan Huang, Zilong Zheng, Jianzhu Ma, Yixin Zhu

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.13786 2023-11-01 cs.CV cs.AI cs.LG 57%

Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Viorica Pătrăucean, Lucas Smaira, Ankush Gupta, Adrià Recasens Continente, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Joseph Heyward, Mateusz Malinowski, Yi Yang, Carl Doersch, Tatiana Matejovicova, Yury Sulsky, Antoine Miech, Alex Frechette, Hanna Klimczak, Raphael Koster, Junlin Zhang, Stephanie Winkler, Yusuf Aytar, Simon Osindero, Dima Damen, Andrew Zisserman, João Carreira

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments 37th Conference on Neural Information Processing Systems (NeurIPS 2023) Track on Datasets and Benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05010 2023-10-10 cs.CV 57%

Building an Open-Vocabulary Video CLIP Model with Better Architectures, Optimization and Data

Zuxuan Wu, Zejia Weng, Wujian Peng, Xitong Yang, Ang Li, Larry S. Davis, Yu-Gang Jiang

专题命中 视频数据与评测 :text-to-video(abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2302.00624

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09558 2023-09-26 cs.CV 57%

ReLER@ZJU Submission to the Ego4D Moment Queries Challenge 2022

Jiayi Shao, Xiaohan Wang, Yi Yang

专题命中 视频数据与评测 :long video(abstract);分类 cs.CV

Comments Accepted to ECCV 2022 Ego4D Workshop; 3rd place in Ego4D Moment Query Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.15055 2023-07-28 cs.CV 57%

PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point Tracking

Yang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein, Leonidas J. Guibas

专题命中 视频数据与评测 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.05721 2023-07-13 cs.CV 57%

HA-ViD: A Human Assembly Video Dataset for Comprehensive Assembly Knowledge Understanding

Hao Zheng, Regina Lee, Yuqian Lu

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02257 2023-04-27 cs.CV 57%

Efficient Annotation and Learning for 3D Hand Pose Estimation: A Survey

Takehiko Ohkawa, Ryosuke Furuta, Yoichi Sato

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06301 2023-04-10 cs.CV 57%

Egocentric Video Task Translation

Zihui Xue, Yale Song, Kristen Grauman, Lorenzo Torresani

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments Accepted by CVPR 2023 (Highlight), Project website: https://vision.cs.utexas.edu/projects/egot2/

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.15616 2023-04-03 cs.CV 57%

Fine-grained Audible Video Description

Xuyang Shen, Dong Li, Jinxing Zhou, Zhen Qin, Bowen He, Xiaodong Han, Aixuan Li, Yuchao Dai, Lingpeng Kong, Meng Wang, Yu Qiao, Yiran Zhong

专题命中 视频数据与评测 :video generation(abstract);分类 cs.CV

Comments accepted to CVPR 2023, Xuyang Shen, Dong Li and Jinxing Zhou contribute equally, code link: github.com/OpenNLPLab/FAVDBench, dataset link: www.avlbench.opennlplab.cn

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08120 2023-03-15 cs.CV 57%

Blind Video Deflickering by Neural Filtering with a Flawed Atlas

Chenyang Lei, Xuanchi Ren, Zhaoxiang Zhang, Qifeng Chen

专题命中 视频数据与评测 :video generation(abstract);分类 cs.CV

Comments To appear in CVPR2023. Code: github.com/ChenyangLEI/All-In-One-Deflicker Website: chenyanglei.github.io/deflicker

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.07210 2022-11-15 cs.CV cs.AI 57%

Grafting Pre-trained Models for Multimodal Headline Generation

Lingfeng Qiao, Chen Wu, Ye Liu, Haoyuan Peng, Di Yin, Bo Ren

专题命中 视频数据与评测 :video-language(abstract);分类 cs.CV

Comments Accepted by EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16644 2022-11-01 cs.CV 57%

Unsupervised Audio-Visual Lecture Segmentation

Darshan Singh S, Anchit Gupta, C. V. Jawahar, Makarand Tapaswi

专题命中 视频数据与评测 :video-language(abstract);分类 cs.CV

Comments 17 pages, 14 figures, 14 tables, Accepted to WACV 2023. Project page: https://cvit.iiit.ac.in/research/projects/cvit-projects/avlectures

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.11107 2022-10-21 cs.CV 57%

FineAction: A Fine-Grained Video Dataset for Temporal Action Localization

Yi Liu, Limin Wang, Yali Wang, Xiao Ma, Yu Qiao

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments Accepted by IEEE T-IP. HomePage: https://deeperaction.github.io/datasets/fineaction

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08389 2022-10-18 cs.CV 57%

Semantic Video Moments Retrieval at Scale: A New Task and a Baseline

Na Li

专题命中 视频数据与评测 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12694 2022-09-27 cs.CV cs.AI 57%

Multi-modal Video Chapter Generation

Xiao Cao, Zitan Chen, Canyu Le, Lei Meng

专题命中 视频数据与评测 :long video(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06513 2022-08-23 cs.CV 57%

Benchmarking the Robustness of Spatial-Temporal Models Against Corruptions

Chenyu Yi, Siyuan Yang, Haoliang Li, Yap-peng Tan, Alex Kot

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments Accepted to NeurIPs 2021 Dataset and Benchmark Track. Our codes are available on https://github.com/Newbeeyoung/Video-Corruption-Robustness

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.14221 2022-08-02 cs.CV 57%

How Severe is Benchmark-Sensitivity in Video Self-Supervised Learning?

Fida Mohammad Thoker, Hazel Doughty, Piyush Bagad, Cees Snoek

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments Accepted in ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12393 2022-07-26 cs.CV 57%

CelebV-HQ: A Large-Scale Video Facial Attributes Dataset

Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, Chen Change Loy

专题命中 视频数据与评测 :video generation(abstract);分类 cs.CV

Comments ECCV 2022. Project Page: https://celebv-hq.github.io/ ; Dataset: https://github.com/CelebV-HQ/CelebV-HQ

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.01558 2022-05-18 cs.CV 57%

Occluded Video Instance Segmentation: A Benchmark

Jiyang Qi, Yan Gao, Yao Hu, Xinggang Wang, Xiaoyu Liu, Xiang Bai, Serge Belongie, Alan Yuille, Philip H. S. Torr, Song Bai

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments IJCV 2022. Project page at https://songbai.site/ovis

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01089 2022-05-03 cs.CV cs.AI cs.LG cs.RO 57%

ComPhy: Compositional Physical Reasoning of Objects and Events from Videos

Zhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding, Antonio Torralba, Joshua B. Tenenbaum, Chuang Gan

专题命中 视频数据与评测 :video reasoning(abstract);分类 cs.CV

Comments ICLR 2022. Project page: https://comphyreasoning.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.10160 2022-04-27 cs.CV 57%

A Multi-Person Video Dataset Annotation Method of Spatio-Temporally Actions

Fan Yang

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.02968 2022-04-07 cs.CV 57%

Temporal Alignment Networks for Long-term Video

Tengda Han, Weidi Xie, Andrew Zisserman

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments CVPR2022 Oral, 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.14821 2022-04-05 cs.CV cs.CL cs.LG 57%

End-to-End Referring Video Object Segmentation with Multimodal Transformers

Adam Botach, Evgenii Zheltonozhskii, Chaim Baskin

专题命中 视频数据与评测 :video understanding(abstract);分类 cs.CV

Comments Accepted to CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00431 2022-03-29 cs.CV cs.AI 57%

MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions

Mattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron, Chen Zhao, Silvio Giancola, Bernard Ghanem

专题命中 视频数据与评测 :video-language(abstract);分类 cs.CV

Comments 12 Pages, 6 Figures, 7 Tables

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.11859 2021-12-08 cs.CV 57%

STEP: Segmenting and Tracking Every Pixel

Mark Weber, Jun Xie, Maxwell Collins, Yukun Zhu, Paul Voigtlaender, Hartwig Adam, Bradley Green, Andreas Geiger, Bastian Leibe, Daniel Cremers, Aljoša Ošep, Laura Leal-Taixé, Liang-Chieh Chen

专题命中 视频数据与评测 :long video(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2021 Track on Datasets and Benchmarks. Code: https://github.com/google-research/deeplab2

详情

展开后加载摘要…

URL PDF HTML 收藏