arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2104.09411 2021-04-20 cs.CV cs.MM 81%

Understanding Chinese Video and Language via Contrastive Multimodal Pre-Training

Chenyi Lei, Shixian Luo, Yong Liu, Wanggui He, Jiamang Wang, Guoxin Wang, Haihong Tang, Chunyan Miao, Houqiang Li

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.04853 2021-03-26 cs.CV cs.AI 81%

Social-STAGE: Spatio-Temporal Multi-Modal Future Trajectory Forecast

Srikanth Malla, Chiho Choi, Behzad Dariush

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments ICRA 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.11624 2021-03-25 cs.CV cs.AI 81%

Multimodal Motion Prediction with Stacked Transformers

Yicheng Liu, Jinghuai Zhang, Liangji Fang, Qinhong Jiang, Bolei Zhou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments CVPR2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.11704 2021-01-29 cs.LG cs.MM cs.SD eess.AS eess.IV 81%

A Case Study of Deep Learning Based Multi-Modal Methods for Predicting the Age-Suitability Rating of Movie Trailers

Mahsa Shafaei, Christos Smailis, Ioannis A. Kakadiaris, Thamar Solorio

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.00536 2021-01-06 cs.CV cs.AI 81%

A Multi-modal Machine Learning Approach and Toolkit to Automate Recognition of Early Stages of Dementia among British Sign Language Users

Xing Liang, Anastassia Angelopoulou, Epaminondas Kapetanios, Bencie Woll, Reda Al-batat, Tyron Woolfe

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Journal ref ECCV 2020 Workshops. Lecture Notes in Computer Science, Vol 12536. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.00073 2021-01-05 cs.CV cs.AI 81%

A Multi-modal Deep Learning Model for Video Thumbnail Selection

Zhifeng Yu, Nanchun Shi

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.10019 2021-01-05 cs.CV cs.AI 81%

Hierarchical Conditional Relation Networks for Multimodal Video Question Answering

Thao Minh Le, Vuong Le, Svetha Venkatesh, Truyen Tran

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Major extension of our CVPR'20 paper to handle long video with text. arXiv admin note: substantial text overlap with arXiv:2002.10698

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.09290 2021-01-01 cs.CV cs.MM 81%

Frame Aggregation and Multi-Modal Fusion Framework for Video-Based Person Recognition

Fangtao Li, Wenzhe Wang, Zihe Liu, Haoran Wang, Chenghao Yan, Bin Wu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments Accepted by MMM 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.09046 2020-11-25 cs.CV cs.CL 81%

A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus

Bowen Zhang, Hexiang Hu, Joonseok Lee, Ming Zhao, Sheide Chammas, Vihan Jain, Eugene Ie, Fei Sha

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.05406 2020-10-13 cs.CL cs.CV 81%

VMSMO: Learning to Generate Multimodal Summary for Video-based News Articles

Mingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan, Dongyan Zhao, Rui Yan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by The 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.09748 2020-08-25 cs.CV cs.AI cs.LG eess.IV eess.SP 81%

Multidomain Multimodal Fusion For Human Action Recognition Using Inertial Sensors

Zeeshan Ahmad, Naimul Khan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.02678 2020-04-29 cs.CV cs.MM 81%

A Local-to-Global Approach to Multi-modal Movie Scene Segmentation

Anyi Rao, Linning Xu, Yu Xiong, Guodong Xu, Qingqiu Huang, Bolei Zhou, Dahua Lin

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Comments CVPR2020. Project page: https://anyirao.com/projects/SceneSeg.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.02205 2020-04-07 cs.CV cs.LG cs.MM 81%

Deep Multimodal Feature Encoding for Video Ordering

Vivek Sharma, Makarand Tapaswi, Rainer Stiefelhagen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments IEEE International Conference on Computer Vision (ICCV) Workshop on Large Scale Holistic Video Understanding. The datasets and code are available at https://github.com/vivoutlaw/tcbp

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.08854 2019-11-21 cs.CV cs.MM 81%

The dynamics of the stomatognathic system from 4D multimodal data

Agnieszka A. Tomaka, Leszek Luchowski, Dariusz Pojda, Michał Tarnawski, Krzysztof Domino

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Chapter 3 in A.Gadomski (ed.): Multiscale Locomotion: Its Active-Matter Addressing Physical Principles; UTP University of Science & Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.02932 2019-10-08 cs.CV cs.IR cs.MM 81%

Multi-Modal Machine Learning for Flood Detection in News, Social Media and Satellite Sequences

Kashif Ahmad, Konstantin Pogorelov, Mohib Ullah, Michael Riegler, Nicola Conci, Johannes Langguth, Ala Al-Fuqaha

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

Journal ref MediaEval 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.02511 2019-03-07 cs.CV cs.AI cs.LG 81%

Learning multimodal representations for sample-efficient recognition of human actions

Miguel Vasco, Francisco S. Melo, David Martins de Matos, Ana Paiva, Tetsunari Inamura

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 7 pages, 6 figures, submitted to 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.12563 2018-12-03 cs.CV cs.MM 81%

Deep Multimodal Learning: An Effective Method for Video Classification

Tianqi Zhao

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.03931 2017-12-12 cs.LG cs.AI cs.CV cs.GR cs.RO 81%

MINOS: Multimodal Indoor Simulator for Navigation in Complex Environments

Manolis Savva, Angel X. Chang, Alexey Dosovitskiy, Thomas Funkhouser, Vladlen Koltun

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments MINOS is a simulator designed to support research on end-to-end navigation

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.05861 2017-10-25 cs.CV cs.LG cs.MM 81%

Continuous Multimodal Emotion Recognition Approach for AVEC 2017

Narotam Singh, Nittin Singh, Abhinav Dhall

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 4 pages, 3 figures, arXiv:1605.06778, arXiv:1512.03385

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.07200 2017-09-22 cs.CV cs.LG cs.MM 81%

Temporal Multimodal Fusion for Video Emotion Classification in the Wild

Valentin Vielzeuf, Stéphane Pateux, Frédéric Jurie

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Journal ref ACM - ICMI 2017, Nov 2017, Glasgow, United Kingdom

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.04508 2017-06-15 cs.MM cs.CV 81%

Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video Classification

Yu-Gang Jiang, Zuxuan Wu, Jinhui Tang, Zechao Li, Xiangyang Xue, Shih-Fu Chang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.07475 2017-02-27 cs.RO cs.AI cs.CV 81%

Sequence-based Multimodal Apprenticeship Learning For Robot Perception and Decision Making

Fei Han, Xue Yang, Yu Zhang, Hao Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 8 pages, 6 figures, accepted by ICRA'17

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.06603 2016-01-26 cs.MM cs.CV 81%

Egocentric Activity Recognition with Multimodal Fisher Vector

Sibo Song, Ngai-Man Cheung, Vijay Chandrasekhar, Bappaditya Mandal, Jie Lin

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 5 pages, 4 figures, ICASSP 2016 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.00599 2016-01-05 cs.CV cs.IR cs.MM 81%

Multimodal Classification of Events in Social Media

Matthias Zeppelzauer, Daniel Schopfhauser

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments Preprint of accepted manuscript for the Elsevier Image and Vision Computing Journal (IMAVIS). The paper will be published by IMAVIS under DOI 10.1016/j.imavis.2015.12.004

详情

展开后加载摘要…

URL PDF HTML 收藏
1511.04024 2015-12-01 cs.CL cs.CV 81%

Multimodal Skip-gram Using Convolutional Pseudowords

Zachary Seymour, Yingming Li, Zhongfei Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1507.04831 2015-07-20 cs.CV cs.LG cs.MM cs.SD 81%

Deep Multimodal Speaker Naming

Yongtao Hu, Jimmy Ren, Jingwen Dai, Chang Yuan, Li Xu, Wenping Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.15349 2024-04-25 eess.SP cs.LG cs.MM 80%

A Survey on Multimodal Wearable Sensor-based Human Action Recognition

Jianyuan Ni, Hao Tang, Syed Tousiful Haque, Yan Yan, Anne H. H. Ngu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments Multimodal Survey for Wearable Sensor-based Human Action Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.05430 2023-09-26 cs.CV 80%

Ensemble Modeling for Multimodal Visual Action Recognition

Jyoti Kini, Sarah Fleischer, Ishan Dave, Mubarak Shah

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 22nd International Conference on Image Analysis and Processing Workshops - Multimodal Action Recognition on the MECCANO Dataset, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07483 2023-07-19 cs.CV 80%

Multimodal Distillation for Egocentric Action Recognition

Gorjan Radevski, Dusan Grujicic, Marie-Francine Moens, Matthew Blaschko, Tinne Tuytelaars

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2023; Codebase released at https://github.com/gorjanradevski/multimodal-distillation

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04916 2023-07-12 cs.CV eess.IV 80%

Rapid Deforestation and Burned Area Detection using Deep Multimodal Learning on Satellite Imagery

Gabor Fodor, Marcos V. Conde

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments CVPR 2023 Workshop on Multimodal Learning for Earth and Environment (MultiEarth)

详情

展开后加载摘要…

URL PDF HTML 收藏