arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2209.06208 2022-10-03 cs.LG cs.AI cs.HC eess.SP 79%

Identification of Cognitive Workload during Surgical Tasks with Multimodal Deep Learning

Kaizhe Jin, Adrian Rubio-Solis, Ravi Naik, Tochukwu Onyeogulu, Amirul Islam, Salman Khan, Izzeddin Teeti, James Kinross, Daniel R Leff, Fabio Cuzzolin, George Mylonas

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.10170 2022-09-22 cs.CV 79%

FV2ES: A Fully End2End Multimodal System for Fast Yet Effective Video Emotion Recognition Inference

Qinglan Wei, Xuling Huang, Yuan Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02080 2022-08-04 cs.CV 79%

A Feature-space Multimodal Data Augmentation Technique for Text-video Retrieval

Alex Falcon, Giuseppe Serra, Oswald Lanz

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted for presentation at 30th ACM International Conference on Multimedia (ACM MM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.01954 2022-08-04 cs.CV 79%

Dilated Context Integrated Network with Cross-Modal Consensus for Temporal Emotion Localization in Videos

Juncheng Li, Junlin Xie, Linchao Zhu, Long Qian, Siliang Tang, Wenqiao Zhang, Haochen Shi, Shengyu Zhang, Longhui Wei, Qi Tian, Yueting Zhuang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ACM Multimedia 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.10123 2022-07-22 cs.CV 79%

Animation from Blur: Multi-modal Blur Decomposition with Motion Guidance

Zhihang Zhong, Xiao Sun, Zhirong Wu, Yinqiang Zheng, Stephen Lin, Imari Sato

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments ECCV2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.05759 2022-07-19 cs.CV 79%

Multimodal Transformer with Variable-length Memory for Vision-and-Language Navigation

Chuang Lin, Yi Jiang, Jianfei Cai, Lizhen Qu, Gholamreza Haffari, Zehuan Yuan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.06187 2022-07-12 cs.CV 79%

Calibrating Class Weights with Multi-Modal Information for Partial Video Domain Adaptation

Xiyu Wang, Yuecong Xu, Kezhi Mao, Jianfei Yang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ACM Multimedia (ACMMM) 2022, update to camera-ready version. 8 pages of text, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.03317 2022-07-08 cs.AI 79%

Multimodal Feature Extraction for Memes Sentiment Classification

Sofiane Ouaari, Tsegaye Misikir Tashu, Tomas Horvath

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02756 2022-07-07 cs.CV 79%

STVGFormer: Spatio-Temporal Video Grounding with Static-Dynamic Cross-Modal Understanding

Zihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin, Tiancai Ye, Wei-Shi Zheng

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Technical report. The 1st place solution in the HC-STVG track of the 4th Person in Context Challenge(2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01241 2022-07-05 cs.CV 79%

OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and Classification

Ye Liu, Lingfeng Qiao, Di Yin, Zhuoxuan Jiang, Xinghua Jiang, Deqiang Jiang, Bo Ren

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by ACM MM 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09852 2022-06-22 cs.CV 79%

M&M Mix: A Multimodal Multiview Transformer Ensemble

Xuehan Xiong, Anurag Arnab, Arsha Nagrani, Cordelia Schmid

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Technical report for Epic-Kitchens challenge 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.02353 2022-06-09 cs.LG cs.CV 79%

Beyond Just Vision: A Review on Self-Supervised Representation Learning on Multimodal and Temporal Data

Shohreh Deldari, Hao Xue, Aaqib Saeed, Jiayuan He, Daniel V. Smith, Flora D. Salim

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 36 pages, 5 figures, 9 tables, Survey paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.01657 2022-05-04 cs.CV 79%

Cross-modal Representation Learning for Zero-shot Action Recognition

Chung-Ching Lin, Kevin Lin, Linjie Li, Lijuan Wang, Zicheng Liu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15086 2022-03-30 cs.CV 79%

X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval

Satya Krishna Gorti, Noel Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, Guangwei Yu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13301 2022-03-29 cs.CV 79%

Multi-modal Multi-label Facial Action Unit Detection with Transformer

Lingfeng Wang, Shisen Wang, Jin Qi

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.12745 2022-03-29 cs.CV 79%

UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection

Ye Liu, Siyuan Li, Yang Wu, Chang Wen Chen, Ying Shan, Xiaohu Qie

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.01122 2022-03-22 cs.CV 79%

M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers

Tsu-Jui Fu, Xin Eric Wang, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.09613 2022-03-18 cs.CV eess.IV 79%

SEN12MS-CR-TS: A Remote Sensing Data Set for Multi-modal Multi-temporal Cloud Removal

Patrick Ebel, Yajin Xu, Michael Schmitt, Xiaoxiang Zhu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Journal ref IEEE Transactions on Geoscience and Remote Sensing, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.09124 2022-02-21 eess.AS cs.SD 79%

Multi-view and Multi-modal Event Detection Utilizing Transformer-based Multi-sensor fusion

Masahiro Yasuda, Yasunori Ohishi, Shoichiro Saito, Noboru Harada

专题命中 视频多模态 :multi-modal(title,abstract);分类 eess.AS

Comments 5 pages, 5 figures, to appear in IEEE ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.03283 2022-02-08 cs.CV 79%

CZU-MHAD: A multimodal dataset for human action recognition utilizing a depth camera and 10 wearable inertial sensors

Xin Chao, Zhenjie Hou, Yujian Mo

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.12713 2022-01-25 cs.CV 79%

Spatio-Contextual Deep Network Based Multimodal Pedestrian Detection For Autonomous Driving

Kinjal Dasgupta, Arindam Das, Sudip Das, Ujjwal Bhattacharya, Senthil Yogamani

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments To be published at IEEE Transactions on Intelligent Transportation Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08443 2021-12-17 cs.LG cs.AI 79%

Event-Aware Multimodal Mobility Nowcasting

Zhaonan Wang, Renhe Jiang, Hao Xue, Flora D. Salim, Xuan Song, Ryosuke Shibasaki

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.07558 2021-12-15 cs.CV eess.IV 79%

Multi-Modal Temporal Attention Models for Crop Mapping from Satellite Time Series

Vivien Sainte Fare Garnot, Loic Landrieu, Nesrine Chehata

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.05379 2021-12-13 cs.CV cs.CR cs.LG 79%

Cross-Modal Transferable Adversarial Attacks from Images to Videos

Zhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang Jiang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.01677 2021-11-03 cs.CV cs.LG 79%

Top1 Solution of QQ Browser 2021 Ai Algorithm Competition Track 1 : Multimodal Video Similarity

Zhuoran Ma, Majing Lou, Xuan Ouyang

专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.00865 2021-11-02 cs.CV eess.IV 79%

MEmoBERT: Pre-training Model with Prompt-based Learning for Multimodal Emotion Recognition

Jinming Zhao, Ruichen Li, Qin Jin, Xinchao Wang, Haizhou Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 4 papges, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.05624 2021-11-01 cs.CV 79%

End-to-end Multi-modal Video Temporal Grounding

Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted in NeurIPS 2021. Project page at https://github.com/wenz116/DRFT

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08270 2021-10-20 cs.LG cs.CL 79%

From Multimodal to Unimodal Attention in Transformers using Knowledge Distillation

Dhruv Agarwal, Tanay Agrawal, Laura M. Ferrari, François Bremond

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Preprint. Final paper accepted at the 17th IEEE International Conference on Advanced Video and Signal-based Surveillance, AVSS 2021, Virtual, November 16-19, 2021. 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08814 2021-10-19 cs.CV 79%

TEAM-Net: Multi-modal Learning for Video Action Recognition with Partial Decoding

Zhengwei Wang, Qi She, Aljosa Smolic

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments To appear in BMVC 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.12671 2021-10-18 cs.CV 79%

Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie Boggust, Rameswar Panda, Brian Kingsbury, Rogerio Feris, David Harwath, James Glass, Michael Picheny, Shih-Fu Chang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments To be presented at ICCV 2021

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 8012-8021

详情

展开后加载摘要…

URL PDF HTML 收藏