arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2009.05702 2020-09-15 cs.RO cs.AI cs.LG cs.SY eess.SY 79%

Risk-Sensitive Sequential Action Control with Multi-Modal Human Trajectory Forecasting for Safe Crowd-Robot Interaction

Haruki Nishimura, Boris Ivanovic, Adrien Gaidon, Marco Pavone, Mac Schwager

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

Comments To appear in 2020 IEEE/RSJ IROS

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.07935 2020-08-25 cs.CV 79%

Describing Unseen Videos via Multi-Modal Cooperative Dialog Agents

Ye Zhu, Yu Wu, Yi Yang, Yan Yan

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments ECCV2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.01148 2020-08-17 cs.RO cs.CV 79%

HAMLET: A Hierarchical Multimodal Attention-based Human Activity Recognition Algorithm

Md Mofijul Islam, Tariq Iqbal

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments To be published in the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2020

Journal ref IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02036 2020-07-07 cs.CV 79%

Modality Shifting Attention Network for Multi-modal Video Question Answering

Junyeong Kim, Minuk Ma, Trung Pham, Kyungsu Kim, Chang D. Yoo

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

Comments CVPR2020 accepted; poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.03844 2020-06-09 cs.CV cs.LG stat.ML 79%

Exploiting Temporal Coherence for Multi-modal Video Categorization

Palash Goyal, Saurabh Sahu, Shalini Ghosh, Chul Lee

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.00654 2020-06-02 cs.LG cs.CV stat.ML 79%

A multimodal approach for multi-label movie genre classification

Rafael B. Mangolin, Rodolfo M. Pereira, Alceu S. Britto, Carlos N. Silla, Valéria D. Feltrim, Diego Bertolini, Yandre M. G. Costa

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 21 pages and 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.13876 2020-05-29 cs.MM 79%

Investigating Correlations of Automatically Extracted Multimodal Features and Lecture Video Quality

Jianwei Shi, Christian Otto, Anett Hoppe, Peter Holtz, Ralph Ewerth

专题命中 视频多模态 :multimodal(title);cross-modal(abstract);分类 cs.MM

Journal ref SALMM '19: Proceedings of the 1st International Workshop on Search as Learning with Multimedia Information, co-located with ACM Multimedia 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.09606 2020-05-20 cs.CL 79%

A Recipe for Creating Multimodal Aligned Datasets for Sequential Tasks

Angela S. Lin, Sudha Rao, Asli Celikyilmaz, Elnaz Nouri, Chris Brockett, Debadeepta Dey, Bill Dolan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments This paper has been accepted to be published at ACL 2020

Journal ref Association of Computational Linguistics 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.08637 2020-05-20 cs.HC cs.CV 79%

Building BROOK: A Multi-modal and Facial Video Database for Human-Vehicle Interaction Research

Xiangjun Peng, Zhentao Huang, Xu Sun

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Conference: ACM CHI Conference on Human Factors in Computing Systems Workshops (CHI'20 Workshops)At: Honolulu, Hawaii, USA URL:https://emergentdatatrails.com

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.06502 2020-04-15 cs.CV cs.LG eess.IV 79%

Unsupervised Multimodal Video-to-Video Translation via Self-Supervised Learning

Kangning Liu, Shuhang Gu, Andres Romero, Radu Timofte

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.12681 2020-04-06 cs.CV cs.LG 79%

What Makes Training Multi-Modal Classification Networks Hard?

Weiyao Wang, Du Tran, Matt Feiszli

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.09691 2020-03-20 cs.CV 79%

Multi-Modal Domain Adaptation for Fine-Grained Action Recognition

Jonathan Munro, Dima Damen

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2020 for an oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.09844 2020-03-13 cs.CV 79%

Adversarial Multimodal Network for Movie Question Answering

Zhaoquan Yuan, Siyuan Sun, Lixin Duan, Xiao Wu, Changsheng Xu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments We will revise the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.06993 2020-03-10 cs.CV cs.RO 79%

Learning Visuomotor Policies for Aerial Navigation Using Cross-Modal Representations

Rogerio Bonatti, Ratnesh Madaan, Vibhav Vineet, Sebastian Scherer, Ashish Kapoor

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.01043 2020-03-03 cs.CL cs.LG stat.ML 79%

Gated Mechanism for Attention Based Multimodal Sentiment Analysis

Ayush Kumar, Jithendra Vepa

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to appear in ICASSP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.12602 2020-02-27 cs.CV 79%

EV-Action: Electromyography-Vision Multi-Modal Action Dataset

Lichen Wang, Bin Sun, Joseph Robinson, Taotao Jing, Yun Fu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments IEEE International Conference on Automatic Face & Gesture Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.01166 2020-02-26 cs.CL 79%

Multimodal Transformer Networks for End-to-End Video-Grounded Dialogue Systems

Hung Le, Doyen Sahoo, Nancy F. Chen, Steven C. H. Hoi

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at ACL 2019 (Long Paper)

Journal ref Association for Computational Linguistics (2019) 5612-5623

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.11657 2020-02-03 cs.CV 79%

Modality Compensation Network: Cross-Modal Adaptation for Action Recognition

Sijie Song, Jiaying Liu, Yanghao Li, Zongming Guo

专题命中 视频多模态 :cross-modal(title);multi-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Trans. on Image Processing, 2020. Project page: http://39.96.165.147/Projects/MCN_tip2020_ssj/MCN_tip_2020_ssj.html

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.00628 2020-01-22 cs.CV cs.LG 79%

Sensor Fusion: Gated Recurrent Fusion to Learn Driving Behavior from Temporal Multimodal Data

Athma Narayanan, Avinash Siravuru, Behzad Dariush

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to Robotics and Automation Letters 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.10982 2019-12-24 cs.CV 79%

DMCL: Distillation Multiple Choice Learning for Multimodal Action Recognition

Nuno C. Garcia, Sarah Adel Bargal, Vitaly Ablavsky, Pietro Morerio, Vittorio Murino, Stan Sclaroff

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.08291 2019-12-17 cs.CV 79%

Correlation Net: Spatiotemporal multimodal deep learning for action recognition

Novanto Yudistira, Takio Kurita

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref Signal Processing: Image Communication, Volume 82, March 2020, 115731

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.09826 2019-11-25 cs.LG cs.CL stat.ML 79%

Factorized Multimodal Transformer for Multimodal Sequential Learning

Amir Zadeh, Chengfeng Mao, Kelly Shi, Yiwei Zhang, Paul Pu Liang, Soujanya Poria, Louis-Philippe Morency

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.03974 2019-11-12 cs.MM cs.IR 79%

A Multimodal CNN-based Tool to Censure Inappropriate Video Scenes

Pedro V. A. de Freitas, Paulo R. C. Mendes, Gabriel N. P. dos Santos, Antonio José G. Busson, Álan Livio Guedes, Sérgio Colcher, Ruy Luiz Milidiú

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.00381 2019-11-04 cs.CV 79%

Multimodal Video-based Apparent Personality Recognition Using Long Short-Term Memory and Convolutional Neural Networks

Süleyman Aslan, Uğur Güdükbay

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.11482 2019-10-28 cs.LG cs.CV stat.ML 79%

Human Action Recognition Using Deep Multilevel Multimodal (M2) Fusion of Depth and Inertial Sensors

Zeeshan Ahmad, Naimul Khan

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.04641 2019-10-11 cs.CV 79%

Cross-modal knowledge distillation for action recognition

Fida Mohammad Thoker, Juergen Gall

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Published in: 2019 IEEE International Conference on Image Processing (ICIP)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.11604 2019-09-26 cs.AI 79%

An Extensible and Personalizable Multi-Modal Trip Planner

Xudong Liu, Christian Fritz, Matthew Klenk

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Published in the Proceedings of the 32nd International Florida Artificial Intelligence Research Society Conference, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.01763 2019-09-05 cs.CV cs.LG 79%

Video Affective Effects Prediction with Multi-modal Fusion and Shot-Long Temporal Context

Jie Zhang, Yin Zhao, Longjun Cai, Chaoping Tu, Wu Wei

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.08498 2019-08-23 cs.CV 79%

EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action Recognition

Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima Damen

专题命中 视频多模态 :audio-visual(title);multi-modal(abstract);分类 cs.CV

Comments Accepted for presentation at ICCV 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.04955 2019-08-15 cs.RO cs.CV cs.HC cs.LG 79%

Probabilistic Multimodal Modeling for Human-Robot Interaction Tasks

Joseph Campbell, Simon Stepputtis, Heni Ben Amor

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Project website: http://interactive-robotics.engineering.asu.edu/interaction-primitives Accompanying video: https://youtu.be/r5AqfxTDfLA

详情

展开后加载摘要…

URL PDF HTML 收藏