arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4749 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4749 篇

2110.06058 2021-10-13 cs.CV 79%

Multi-Modal Interaction Graph Convolutional Network for Temporal Language Localization in Videos

Zongmeng Zhang, Xianjing Han, Xuemeng Song, Yan Yan, Liqiang Nie

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing

Journal ref in IEEE Transactions on Image Processing, vol. 30, pp. 8265-8277, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.13666 2021-09-29 cs.CV cs.RO 79%

Fail-Safe Human Detection for Drones Using a Multi-Modal Curriculum Learning Approach

Ali Safa, Tim Verbelen, Ilja Ocket, André Bourdoux, Francky Catthoor, Georges G. E. Gielen

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.04735 2021-09-13 cs.CV 79%

Temporal Pyramid Transformer with Multimodal Interaction for Video Question Answering

Min Peng, Chongyang Wang, Yuan Gao, Yu Shi, Xiang-Dong Zhou

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Submitted to AAAI'22

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.03619 2021-08-10 cs.CV 79%

Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action Detection

Rui Dai, Srijan Das, Francois Bremond

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.12589 2021-07-28 cs.CV 79%

Cross-modal Consensus Network for Weakly Supervised Temporal Action Localization

Fa-Ting Hong, Jia-Chang Feng, Dan Xu, Ying Shan, Wei-Shi Zheng

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments ACM International Conference on Multimedia, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.14435 2021-07-22 cs.CV 79%

DanHAR: Dual Attention Network For Multimodal Human Activity Recognition Using Wearable Sensors

Wenbin Gao, Lei Zhang, Qi Teng, Jun He, Hao Wu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.09504 2021-07-21 cs.CV 79%

Multi-Modal Temporal Convolutional Network for Anticipating Actions in Egocentric Videos

Olga Zatsarynna, Yazan Abu Farha, Juergen Gall

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments CVPR Precognition Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.04187 2021-07-16 cs.CV 79%

A Multi-modal and Multi-task Learning Method for Action Unit and Expression Recognition

Yue Jin, Tianqing Zheng, Chao Gao, Guoqiang Xu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 5 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03009 2021-07-13 cs.CV 79%

Multi-modal Affect Analysis using standardized data within subjects in the Wild

Sachihiro Youoku, Takahisa Yamamoto, Junya Saito, Akiyoshi Uchida, Xiaoyu Mi, Ziqiang Shi, Liu Liu, Zhongling Liu, Osafumi Nakayama, Kentaro Murase

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 6 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.09412 2021-07-07 cs.CV 79%

Searching Multi-Rate and Multi-Modal Temporal Enhanced Networks for Gesture Recognition

Zitong Yu, Benjia Zhou, Jun Wan, Pichao Wang, Haoyu Chen, Xin Liu, Stan Z. Li, Guoying Zhao

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Submitted to IEEE Transactions on Image Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.14137 2021-06-29 cs.CV 79%

Building a Video-and-Language Dataset with Human Actions for Multimodal Logical Inference

Riko Suzuki, Hitomi Yanaka, Koji Mineshima, Daisuke Bekki

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to MMSR I

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08252 2021-05-19 cs.CV 79%

Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal Matching

Bofeng Wu, Guocheng Niu, Jun Yu, Xinyan Xiao, Jian Zhang, Hua Wu

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.08833 2021-05-04 cs.CV 79%

Skeleton Aware Multi-modal Sign Language Recognition

Songyao Jiang, Bin Sun, Lichen Wang, Yue Bai, Kunpeng Li, Yun Fu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments This is a preprint version of our work SAM-SLR that ranked 1st at CVPR2021 Challenge on Large Scale Signer Independent Isolated Sign Language Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10139 2021-04-21 cs.CL 79%

Towards Solving Multimodal Comprehension

Pritish Sahu, Karan Sikka, Ajay Divakaran

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.03086 2021-04-06 eess.IV cs.CV 79%

A Multi-Modal Respiratory Disease Exacerbation Prediction Technique Based on a Spatio-Temporal Machine Learning Architecture

Rohan Tan Bhowmik

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Updated Title, References, and Acknowledgements

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.06747 2021-03-30 cs.CV 79%

ChallenCap: Monocular 3D Capture of Challenging Human Performances using Multi-Modal References

Yannan He, Anqi Pang, Xin Chen, Han Liang, Minye Wu, Yuexin Ma, Lan Xu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.10572 2021-03-23 cs.MM 79%

Quantum-inspired Multimodal Fusion for Video Sentiment Analysis

Qiuchi Li, Dimitris Gkoumas, Christina Lioma, Massimo Melucci

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

Comments Post-print accepted by Information Fusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.09320 2021-02-19 cs.CV 79%

Combining Events and Frames using Recurrent Asynchronous Multimodal Networks for Monocular Depth Prediction

Daniel Gehrig, Michelle Rüegg, Mathias Gehrig, Javier Hidalgo Carrio, Davide Scaramuzza

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref IEEE Robotics and Automation Letters (RA-L), 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.04762 2021-02-10 cs.CV 79%

Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network

Linwei Ye, Mrigank Rochan, Zhi Liu, Xiaoqin Zhang, Yang Wang

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV

Comments 14 pages, 8 figures. arXiv admin note: substantial text overlap with arXiv:1904.04745

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.04727 2021-02-10 cs.CV 79%

Fashion Focus: Multi-modal Retrieval System for Video Commodity Localization in E-commerce

Yanhao Zhang, Qiang Wang, Pan Pan, Yun Zheng, Cheng Da, Siyang Sun, Yinghui Xu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments accepted by AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.08165 2021-01-21 cs.CV 79%

Video Relation Detection with Trajectory-aware Multi-modal Features

Wentao Xie, Guanghui Ren, Si Liu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.07339 2021-01-21 cs.CL 79%

MONAH: Multi-Modal Narratives for Humans to analyze conversations

Joshua Y. Kim, Greyson Y. Kim, Chunfeng Liu, Rafael A. Calvo, Silas C. R. Taylor, Kalina Yacef

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.11851 2020-12-23 cs.CV 79%

Predicting Online Video Advertising Effects with Multimodal Deep Learning

Jun Ikeda, Hiroyuki Seshime, Xueting Wang, Toshihiko Yamasaki

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at International Conference on Pattern Recognition 2020 (ICPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.03186 2020-12-11 cs.CV 79%

Noise Estimation Using Density Estimation for Self-Supervised Multimodal Learning

Elad Amrani, Rami Ben-Ari, Daniel Rotman, Alex Bronstein

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted to AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00514 2020-12-02 cs.CV cs.RO 79%

Multi-Modal Hybrid Architecture for Pedestrian Action Prediction

Amir Rasouli, Tiffany Yau, Mohsen Rohani, Jun Luo

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 7 pages, 4 Figures, 3 tables, submitted to ICRA 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.04417 2020-11-11 cs.CV 79%

Disentangle, align and fuse for multimodal and semi-supervised image segmentation

Agisilaos Chartsias, Giorgos Papanastasiou, Chengjia Wang, Scott Semple, David E. Newby, Rohan Dharmakumar, Sotirios A. Tsaftaris

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref IEEE Transactions on Medical Imaging (2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.13839 2020-10-28 cs.LG cs.CL 79%

VisualHints: A Visual-Lingual Environment for Multimodal Reinforcement Learning

Thomas Carta, Subhajit Chaudhury, Kartik Talamadupula, Michiaki Tatsubori

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

Comments Code is available at http://ibm.biz/VisualHints

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.13626 2020-10-27 cs.CV cs.LG 79%

Classification of Important Segments in Educational Videos using Multimodal Features

Junaid Ahmed Ghauri, Sherzod Hakimov, Ralph Ewerth

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments Proceedings of the CIKM 2020 Workshops, October 19 to 20, Galway, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.11884 2020-10-23 cs.CV cs.HC 79%

AEGIS: A real-time multimodal augmented reality computer vision based system to assist facial expression recognition for individuals with autism spectrum disorder

James Ren Hou Lee, Alexander Wong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

Comments 4 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.08018 2020-09-18 cs.CL cs.IR 79%

Multi-modal Summarization for Video-containing Documents

Xiyan Fu, Jun Wang, Zhenglu Yang

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CL

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏