arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

1905.13540 2019-06-03 cs.CV cs.LG stat.ML 79%

Gaining Extra Supervision via Multi-task learning for Multi-Modal Video Question Answering

Junyeong Kim, Minuk Ma, Kyungsu Kim, Sungjin Kim, Chang D. Yoo

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to IJCNN2019, oral

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.01263 2019-05-21 cs.CV 79%

Multimodal Explanations by Predicting Counterfactuality in Videos

Atsushi Kanehira, Kentaro Takemoto, Sho Inayoshi, Tatsuya Harada

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Camera ready version of CVPR'19

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.01790 2019-05-14 cs.MM 79%

A multimodal lossless coding method for skeletons in videos

Mingzhou Liu, Xiaoyi He, Weiyao Lin, Xintong Han, Yanmin Zhu, Hongtao Lu, Hongkai Xiong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments This manuscript is the accepted version for ICMEW (IEEE Intl. Conf. Multimedia & Expo Workshop), IEEE Intl. Conf. Multimedia & Expo Workshop (ICME), 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.03879 2019-04-30 cs.CV 79%

Cross and Learn: Cross-Modal Self-Supervision

Nawid Sayed, Biagio Brattoli, Björn Ommer

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments GCPR 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.04357 2019-04-10 cs.CV 79%

Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering

Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, Heng Huang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.04927 2018-11-28 cs.CV 79%

VIPL-HR: A Multi-modal Database for Pulse Estimation from Less-constrained Face Video

Xuesong Niu, Hu Han, Shiguang Shan, Xilin Chen

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.04595 2018-11-13 cs.CV 79%

Holistic Multi-modal Memory Network for Movie Question Answering

Anran Wang, Anh Tuan Luu, Chuan-Sheng Foo, Hongyuan Zhu, Yi Tay, Vijay Chandrasekhar

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.10641 2018-07-30 cs.CV 79%

Multimodal Classification with Deep Convolutional-Recurrent Neural Networks for Electroencephalography

Chuanqi Tan, Fuchun Sun, Wenchang Zhang, Jianhua Chen, Chunfang Liu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures

Journal ref Neural Information Processing. 2017:767-776

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.10319 2018-06-28 cs.CV 79%

Exploiting Spatial-Temporal Modelling and Multi-Modal Fusion for Human Action Recognition

Dongliang He, Fu Li, Qijie Zhao, Xiang Long, Yi Fu, Shilei Wen

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.02609 2018-06-08 cs.CV 79%

Learning Multi-Modal Self-Awareness Models for Autonomous Vehicles from Human Driving

Mahdyar Ravanbakhsh, Mohamad Baydoun, Damian Campo, Pablo Marin, David Martin, Lucio Marcenaro, Carlo S. Regazzoni

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments FUSION 2018 - 21st International Conference on Information Fusion, Cambridge, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.11372 2018-05-30 cs.CV 79%

"How to rate a video game?" - A prediction system for video games based on multimodal information

Vishal Batchu, Varshit Battu, Murali Krishna Reddy, Radhika Mamidi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments ICPRAI-18

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.06057 2018-04-18 cs.MM 79%

Multimodal Co-Training for Selecting Good Examples from Webly Labeled Video

Ryota Hinami, Junwei Liang, Shin'ichi Satoh, Alexander Hauptmann

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.05939 2018-04-04 cs.CV q-bio.NC 79%

AJILE Movement Prediction: Multimodal Deep Learning for Natural Human Neural Recordings and Video

Nancy Xin Ru Wang, Ali Farhadi, Rajesh Rao, Bingni Brunton

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref Thirty-Second AAAI Conference On Artificial Intelligence (2018)

详情

展开后加载摘要…

URL PDF HTML 收藏
1703.08513 2018-02-08 cs.CL q-bio.NC 79%

Interactive Natural Language Acquisition in a Multi-modal Recurrent Neural Architecture

Stefan Heinrich, Stefan Wermter

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL

Comments Received 25 June 2016; Accepted 1 February 2017

Journal ref Connection Science, vol 30, No 1, pp. 99-133, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.07721 2018-01-15 cs.CV 79%

An Order Preserving Bilinear Model for Person Detection in Multi-Modal Data

Oytun Ulutan, Benjamin S. Riggan, Nasser M. Nasrabadi, B. S. Manjunath

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.03321 2017-11-28 stat.ML cs.CV q-bio.NC 79%

A deep learning architecture for temporal sleep stage classification using multivariate and multimodal time series

Stanislas Chambon, Mathieu Galtier, Pierrick Arnal, Gilles Wainrib, Alexandre Gramfort

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.08569 2017-11-27 cs.CV 79%

Geometric Cross-Modal Comparison of Heterogeneous Sensor Data

Christopher J. Tralie, Abraham Smith, Nathan Borggren, Jay Hineman, Paul Bendich, Peter Zulch, John Harer

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments 10 pages, 13 figures, Proceedings of IEEE Aeroconf 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.05219 2017-10-17 cs.LG cs.AI 79%

Mental Sampling in Multimodal Representations

Jian-Qiao Zhu, Adam N. Sanborn, Nick Chater

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.05461 2017-07-11 cs.CV 79%

Truly Multi-modal YouTube-8M Video Classification with Video, Audio, and Text

Zhe Wang, Kingsley Kuan, Mathieu Ravaut, Gaurav Manek, Sibo Song, Yuan Fang, Seokhwan Kim, Nancy Chen, Luis Fernando D'Haro, Luu Anh Tuan, Hongyuan Zhu, Zeng Zeng, Ngai Man Cheung, Georgios Piliouras, Jie Lin, Vijay Chandrasekhar

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 8 pages, Accepted to CVPR'17 Workshop on YouTube-8M Large-Scale Video Understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.05103 2017-05-16 cs.MM cs.IR 79%

Generative Adversarial Networks for Multimodal Representation Learning in Video Hyperlinking

Vedran Vukotic, Christian Raymond, Guillaume Gravier

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments 4 pages, 1 figure, 2 tables, published at ACM International Conference in Multimedia Retrieval (ICMR) 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.05244 2017-04-14 cs.CL cs.IR 79%

Select-Additive Learning: Improving Generalization in Multimodal Sentiment Analysis

Haohan Wang, Aaksha Meghawat, Louis-Philippe Morency, Eric P. Xing

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Supplementary files at: http://www.cs.cmu.edu/~haohanw/document/sal_supp.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.01638 2017-02-07 cs.CV 79%

Concurrent Activity Recognition with Multimodal CNN-LSTM Structure

Xinyu Li, Yanyi Zhang, Jianyu Zhang, Shuhong Chen, Ivan Marsic, Richard A. Farneth, Randall S. Burd

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 14 pages, 12 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
1603.07120 2016-12-28 cs.CV 79%

Deep Multimodal Feature Analysis for Action Recognition in RGB+D Videos

Amir Shahroudy, Tian-Tsong Ng, Yihong Gong, Gang Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1412.7006 2016-11-15 cs.CV cs.LG cs.NE 79%

Multi-modal Sensor Registration for Vehicle Perception via Deep Neural Networks

Michael Giering, Vivek Venugopalan, Kishore Reddy

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 7 pages, double column, IEEE format, accepted at IEEE HPEC 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.03112 2016-10-12 cs.CL 79%

Leveraging Recurrent Neural Networks for Multimodal Recognition of Social Norm Violation in Dialog

Tiancheng Zhao, Ran Zhao, Zhao Meng, Justine Cassell

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Submitted to NIPS Workshop. arXiv admin note: text overlap with arXiv:1608.02977 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0611138 2016-08-16 cs.AI 79%

Functional Brain Imaging with Multi-Objective Multi-Modal Evolutionary Optimization

Vojtech Krmicek, Michèle Sebag

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Journal ref Dans PPSN'06, 4193 (2006) 382-391

详情

展开后加载摘要…

URL PDF HTML 收藏
1607.04780 2016-07-19 cs.CV cs.LG 79%

Exploiting Multi-modal Curriculum in Noisy Web Data for Large-scale Concept Learning

Junwei Liang, Lu Jiang, Deyu Meng, Alexander Hauptmann

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1607.02652 2016-07-12 cs.HC cs.CV 79%

Multimodal Affect Recognition using Kinect

Amol Patwardhan, Gerald Knapp

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 9 pages, 2 tables, 1 figure, Peer reviewed in ACM TIST

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.03237 2016-06-13 cs.CV 79%

Survey on RGB, 3D, Thermal, and Multimodal Approaches for Facial Expression Recognition: History, Trends, and Affect-related Applications

Ciprian Corneanu, Marc Oliu, Jeffrey F. Cohn, Sergio Escalera

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.00022 2016-01-06 cs.CV 79%

Event Specific Multimodal Pattern Mining with Image-Caption Pairs

Hongzhi Li, Joseph G. Ellis, Shih-Fu Chang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏