arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4772 篇

2212.00773 2024-02-09 cs.CV 83%

FakeOut: Leveraging Out-of-domain Self-supervision for Multi-modal Video Deepfake Detection

Gil Knafo, Ohad Fried

专题命中 视频多模态 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17797 2024-02-01 cs.CV 83%

M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval

Xingning Dong, Zipeng Feng, Chunluan Zhou, Xuzheng Yu, Ming Yang, Qingpei Guo

专题命中 视频多模态 :multi-modal(title,abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01674 2024-01-04 cs.CV 83%

Transformer RGBT Tracking with Spatio-Temporal Multimodal Tokens

Dengdi Sun, Yajie Pan, Andong Lu, Chenglong Li, Bin Luo

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12721 2023-12-21 cs.CV 83%

Cross-Modal Reasoning with Event Correlation for Video Question Answering

Chengxiang Yin, Zhengping Che, Kun Wu, Zhiyuan Xu, Qinru Qiu, Jian Tang

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.08083 2023-12-14 cs.CV 83%

VCD: Visual Causality Discovery for Cross-Modal Question Reasoning

Yang Liu, Ying Tan, Jingzhou Luo, Weixing Chen

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 12 pages, 6 figures. arXiv admin note: substantial text overlap with arXiv:2207.12647

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15915 2023-09-29 cs.CV 83%

Zero-Shot and Few-Shot Video Question Answering with Multi-Modal Prompts

Deniz Engin, Yannis Avrithis

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments ICCV2023 CLVL Workshop (Oral). Project page: https://engindeniz.github.io/vitis

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15313 2023-09-28 cs.CV 83%

M$^{3}$3D: Learning 3D priors using Multi-Modal Masked Autoencoders for 2D image and video understanding

Muhammad Abdullah Jamal, Omid Mohareri

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.15082 2023-09-27 cs.CV 83%

RPEFlow: Multimodal Fusion of RGB-PointCloud-Event for Joint Optical Flow and Scene Flow Estimation

Zhexiong Wan, Yuxin Mao, Jing Zhang, Yuchao Dai

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ICCV 2023. Project page: https://npucvr.github.io/RPEFlow Code: https://github.com/danqu130/RPEFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.02041 2023-09-06 cs.CV 83%

Learning Cross-Modal Affinity for Referring Video Object Segmentation Targeting Limited Samples

Guanghui Li, Mingqi Gao, Heng Liu, Xiantong Zhen, Feng Zheng

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.12244 2023-08-24 cs.CV 83%

Multimodal Channel-Mixing: Channel and Spatial Masked AutoEncoder on Facial Action Unit Detection

Xiang Zhang, Huiyuan Yang, Taoyue Wang, Xiaotian Li, Lijun Yin

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09775 2023-08-22 cs.CV 83%

Long-range Multimodal Pretraining for Movie Understanding

Dawit Mureja Argaw, Joon-Young Lee, Markus Woodson, In So Kweon, Fabian Caba Heilbron

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14652 2023-06-01 cs.CL 83%

Denoising Bottleneck with Mutual Information Maximization for Video Multimodal Fusion

Shaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu, Binghuai Lin, Yunbo Cao, Zhifang Sui

专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CL

Comments Accept at ACL2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.10842 2023-05-09 cs.CV cs.LG cs.RO 83%

MMRNet: Improving Reliability for Multimodal Object Detection and Segmentation for Bin Picking via Multimodal Redundancy

Yuhao Chen, Hayden Gunraj, E. Zhixuan Zeng, Robbie Meyer, Maximilian Gilles, Alexander Wong

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted to CVPR TCV Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.05419 2023-05-01 cs.CV 83%

Multimodal Graph Learning for Deepfake Detection

Zhiyuan Yan, Peng Sun, Yubo Lang, Shuo Du, Shanzhuo Zhang, Wei Wang, Lei Liu

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.17285 2023-03-31 cs.CV 83%

Decomposed Cross-modal Distillation for RGB-based Temporal Action Detection

Pilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee, Hyeran Byun

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.14905 2023-03-28 cs.CV cs.AI cs.CL cs.LG cs.MM 83%

Multi-Modal Few-Shot Temporal Action Detection

Sauradip Nag, Mengmeng Xu, Xiatian Zhu, Juan-Manuel Perez-Rua, Bernard Ghanem, Yi-Zhe Song, Tao Xiang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10826 2023-03-28 cs.CV 83%

Visual Prompt Multi-Modal Tracking

Jiawen Zhu, Simiao Lai, Xin Chen, Dong Wang, Huchuan Lu

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments Accepted by CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.00040 2023-03-27 cs.CV 83%

Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-Training

Dezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin, Yang Liu

专题命中 视频多模态 :image-text(title,abstract);multi-modal(abstract);分类 cs.CV

Comments CVPR2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.10973 2022-12-05 cs.MM 83%

FakeSV: A Multimodal Benchmark with Rich Social Context for Fake News Detection on Short Video Platforms

Peng Qi, Yuyan Bu, Juan Cao, Wei Ji, Ruihao Shui, Junbin Xiao, Danding Wang, Tat-Seng Chua

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.MM

Comments To appear in AAAI 2023 AISI track. This version contains appendix with additional details

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12833 2022-07-27 cs.CV 83%

Multimodal-GuideNet: Gaze-Probe Bidirectional Guidance in Obstetric Ultrasound Scanning

Qianhui Men, Clare Teng, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble

专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Early accepted by MICCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.03014 2022-03-29 cs.CV 83%

Learnable Irrelevant Modality Dropout for Multimodal Action Recognition on Modality-Specific Annotated Videos

Saghir Alfasly, Jian Lu, Chen Xu, Yuru Zou

专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.09322 2021-11-16 cs.CV 83%

MM-ViT: Multi-Modal Video Transformer for Compressed Video Action Recognition

Jiawei Chen, Chiu Man Ho

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments Winter Conference on Applications of Computer Vision (WACV) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.12868 2021-08-31 cs.CV 83%

A Multimodal Framework for Video Ads Understanding

Zejia Weng, Lingchen Meng, Rui Wang, Zuxuan Wu, Yu-Gang Jiang

专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 4 pages; 2 figures; ACM MM 2021 workshop; Tencent Advertising Algorithm Competition ACM Multimedia 2021 Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.11974 2021-08-30 cs.CV 83%

Learning Cross-modal Contrastive Features for Video Domain Adaptation

Donghyun Kim, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu, Stan Sclaroff, Kate Saenko, Manmohan Chandraker

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted in ICCV'21

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.12465 2021-04-27 cs.CV cs.AI cs.CL cs.MM 83%

GPT2MVS: Generative Pre-trained Transformer-2 for Multi-modal Video Summarization

Jia-Hong Huang, Luka Murn, Marta Mrak, Marcel Worring

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments This paper is accepted by ACM International Conference on Multimedia Retrieval (ICMR), 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.12667 2020-10-27 cs.CV 83%

Self-Supervised Learning by Cross-Modal Audio-Video Clustering

Humam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani, Bernard Ghanem, Du Tran

专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV

Comments Accepted to NeurIPS 2020 (spotlight presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.10639 2020-07-22 cs.CV 83%

Multi-modal Transformer for Video Retrieval

Valentin Gabeur, Chen Sun, Karteek Alahari, Cordelia Schmid

专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

Comments ECCV 2020 (spotlight paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.09944 2019-10-28 cs.CV 83%

Watch, Listen and Tell: Multi-modal Weakly Supervised Dense Event Captioning

Tanzila Rahman, Bicheng Xu, Leonid Sigal

专题命中 视频多模态 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV

Journal ref ICCV2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.03989 2018-10-23 cs.CV 83%

Image-to-Video Person Re-Identification by Reusing Cross-modal Embeddings

Zhongwei Xie, Lin Li, Xian Zhong, Luo Zhong

专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments under review for Pattern Recognition Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.05848 2018-10-01 cs.CV 83%

Towards Good Practices for Multi-modal Fusion in Large-scale Video Classification

Jinlai Liu, Zehuan Yuan, Changhu Wang

专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV

Comments ECCV YouTube-8M workshop general paper

详情

展开后加载摘要…

URL PDF HTML 收藏