arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4585 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4585 篇

2108.04343 2021-08-11 cs.AI cs.LG 79%

Towards a Generic Multimodal Architecture for Batch and Streaming Big Data Integration

Siham Yousfi, Maryem Rhanoui, Dalila Chiadmi

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Journal ref Journal of Computer Science, Volume 15 No. 1, 2019, 207-220

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06170 2021-08-10 cs.CV 79%

ViNet: Pushing the limits of Visual Modality for Audio-Visual Saliency Prediction

Samyak Jain, Pradeep Yarlagadda, Shreyank Jyoti, Shyamgopal Karthik, Ramanathan Subramanian, Vineet Gandhi

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Appearing in the proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2021) (camera-ready version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10557 2021-08-04 cs.IR cs.MM 79%

Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-Overs

Tingtian Li, Zixun Sun, Haoruo Zhang, Jin Li, Ziming Wu, Hui Zhan, Yipeng Yu, Hengcan Shi

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.MM

Comments accepted by ACM SIGIR '21

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.00096 2021-08-03 cs.LG cs.CL 79%

Multi-Modal Detection of Alzheimer's Disease from Speech and Text

Amish Mittal, Sourav Sahoo, Arnhav Datar, Juned Kadiwala, Hrithwik Shalu, Jimson Mathew

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL

Comments 9 pages, 3 figures, Accepted in BIOKDD 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.06592 2021-07-27 eess.AS cs.SD eess.IV 79%

Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection

Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, Haizhou Li

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments ACM Multimedia 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.08855 2021-07-21 cs.NE cs.AI q-bio.NC 79%

A Biologically Plausible Audio-Visual Integration Model for Continual Learning

Wenjie Chen, Fengtong Du, Ye Wang, Lihong Cao

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI

Comments Accepted by 2021 International Joint Conference on Neural Networks

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.07452 2021-07-15 cs.MM cs.NI 79%

Software Defined Media: Virtualization of Audio-Visual Services

Manabu Tsukada, Keiko Ogawa, Masahiro Ikeda, Takuro Sone, Kenta Niwa, Shoichiro Saito, Takashi Kasuya, Hideki Sunahara, Hiroshi Esaki

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.MM

Comments IEEE International Conference on Communications (ICC2017), Paris, France, 21-25 May 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.01571 2021-07-06 cs.CL 79%

Audio-Oriented Multimodal Machine Comprehension: Task, Dataset and Model

Zhiqi Huang, Fenglin Liu, Xian Wu, Shen Ge, Helin Wang, Wei Fan, Yuexian Zou

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments AAAI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.01569 2021-07-06 cs.CL cs.LG 79%

Cross-Modal Transformer-Based Neural Correction Models for Automatic Speech Recognition

Tomohiro Tanaka, Ryo Masumura, Mana Ihori, Akihiko Takashima, Takafumi Moriya, Takanori Ashihara, Shota Orihashi, Naoki Makishima

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CL

Comments Accepted to Interspeech 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.16153 2021-07-01 cs.IR cs.SD eess.AS 79%

Multi-Modal Chorus Recognition for Improving Song Search

Jiaan Wang, Zhixu Li, Binbin Gu, Tingyi Zhang, Qingsheng Liu, Zhigang Chen

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments Accepted at the 30th International Conference on Artificial Neural Networks (ICANN 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.12037 2021-06-24 cs.CV 79%

Listen to Your Favorite Melodies with img2Mxml, Producing MusicXML from Sheet Music Image by Measure-based Multimodal Deep Learning-driven Assembly

Tomoyuki Shishido, Fehmiju Fati, Daisuke Tokushige, Yasuhiro Ono

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments 19 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.08513 2021-06-17 cs.CV 79%

Watching Too Much Television is Good: Self-Supervised Audio-Visual Representation Learning from Movies and TV Shows

Mahdi M. Kalayeh, Nagendra Kamath, Lingyi Liu, Ashok Chandrashekar

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.06840 2021-06-17 cs.SD eess.AS 79%

Deep Learning Frameworks Applied For Audio-Visual Scene Classification

Lam Pham, Alexander Schindler, Mina Schütz, Jasmin Lampert, Sven Schlarb, Ross King

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.06939 2021-06-15 cs.CV 79%

Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning

Shaobo Min, Qi Dai, Hongtao Xie, Chuang Gan, Yongdong Zhang, Jingdong Wang

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.02901 2021-06-15 eess.AS cs.SD 79%

S2VC: A Framework for Any-to-Any Voice Conversion with Self-Supervised Pretrained Representations

Jheng-hao Lin, Yist Y. Lin, Chung-Ming Chien, Hung-yi Lee

专题命中 音频语音多模态 :any-to-any(title,abstract);分类 eess.AS

Comments Accepted by INTERSPEECH 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.03786 2021-06-10 eess.AS cs.LG cs.SD eess.SP 79%

EMA2S: An End-to-End Multimodal Articulatory-to-Speech System

Yu-Wen Chen, Kuo-Hsuan Hung, Shang-Yi Chuang, Jonathan Sherman, Wen-Chin Huang, Xugang Lu, Yu Tsao

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.02948 2021-06-08 q-bio.NC cs.AI cs.LG 79%

Neural dSCA: demixing multimodal interaction among brain areas during naturalistic experiments

Yu Takagi, Laurence T. Hunt, Ryu Ohata, Hiroshi Imamizu, Jun-ichiro Hirayama

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.00639 2021-06-08 eess.AS cs.SD eess.SP 79%

Multi-modal Point-of-Care Diagnostics for COVID-19 Based On Acoustics and Symptoms

Srikanth Raj Chetupalli, Prashant Krishnan, Neeraj Sharma, Ananya Muguli, Rohit Kumar, Viral Nanda, Lancelot Mark Pinto, Prasanta Kumar Ghosh, Sriram Ganapathy

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments The Manuscript is submitted to IEEE-EMBS Journal of Biomedical and Health Informatics on June 1, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.12855 2021-05-28 cs.CV 79%

Multi-Modal Semantic Inconsistency Detection in Social Media News Posts

Scott McCrae, Kehan Wang, Avideh Zakhor

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.12536 2021-05-27 eess.AS cs.SD 79%

Exploiting Temporal Dependencies for Cross-Modal Music Piece Identification

Luis Carvalho, Gerhard Widmer

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 eess.AS

Comments 5 pages, 3 figures

Journal ref Proceedings of the 29th European Signal Processing Conference (EUSIPCO 2021), Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.11087 2021-05-25 cs.CV 79%

Recent Advances and Trends in Multimodal Deep Learning: A Review

Jabeen Summaira, Xi Li, Amin Muhammad Shoib, Songyuan Li, Jabbar Abdul

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.06107 2021-05-14 cs.SD cs.RO eess.AS 79%

Multi-target DoA Estimation with an Audio-visual Fusion Mechanism

Xinyuan Qian, Maulik Madhavi, Zexu Pan, Jiadong Wang, Haizhou Li

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments ICASSP 2021 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.14150 2021-05-04 eess.AS cs.LG 79%

FragmentVC: Any-to-Any Voice Conversion by End-to-End Extracting and Fusing Fine-Grained Voice Fragments With Attention

Yist Y. Lin, Chung-Ming Chien, Jheng-Hao Lin, Hung-yi Lee, Lin-shan Lee

专题命中 音频语音多模态 :any-to-any(title,abstract);分类 eess.AS

Comments To appear in the proceedings of ICASSP 2021, equal contribution from first two authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.14799 2021-05-03 cs.MM 79%

Cross-Modal Music-Video Recommendation: A Study of Design Choices

Laure Pretet, Gael Richard, Geoffroy Peeters

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.12807 2021-04-29 cs.SD eess.AS 79%

Multimodal Self-Supervised Learning of General Audio Representations

Luyu Wang, Pauline Luc, Adria Recasens, Jean-Baptiste Alayrac, Aaron van den Oord

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10715 2021-04-23 cs.LG cs.AI 79%

Uncertainty-Aware Boosted Ensembling in Multi-Modal Settings

Utkarsh Sarawgi, Rishab Khincha, Wazeer Zulfikar, Satrajit Ghosh, Pattie Maes

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

Comments Accepted at IJCNN 2021, to appear in IEEE proceedings. Equal contributions from US, RK and WZ

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.09482 2021-04-20 eess.AS cs.SD 79%

Fusing information streams in end-to-end audio-visual speech recognition

Wentao Yu, Steffen Zeiler, Dorothea Kolossa

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments 5 pages

Journal ref Published in International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.12283 2021-04-13 cs.CL cs.LG 79%

ST-BERT: Cross-modal Language Model Pre-training For End-to-end Spoken Language Understanding

Minjeong Kim, Gyuwan Kim, Sang-Woo Lee, Jung-Woo Ha

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CL

Comments ICASSP 2021; 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.10652 2021-04-08 cs.CL cs.LG 79%

Self-Supervised learning with cross-modal transformers for emotion recognition

Aparna Khare, Srinivas Parthasarathy, Shiva Sundaram

专题命中 音频语音多模态 :cross-modal(title);multi-modal(abstract);分类 cs.CL

Comments To appear in SLT2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.01616 2021-03-03 cs.CL cs.LG 79%

Interpretable Multi-Modal Hate Speech Detection

Prashanth Vijayaraghavan, Hugo Larochelle, Deb Roy

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

Comments 5 pages, Accepted at the International Conference on Machine Learning AI for Social Good Workshop, Long Beach, United States, 2019

Journal ref ICML Workshop on AI for Social Good, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏