arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2112.04748 2022-09-13 cs.SD cs.AI eess.AS 62%

LipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction and Lip Reading

Leyuan Qu, Cornelius Weber, Stefan Wermter

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS

Comments ACCEPTED IN IEEE Transactions on Neural Networks and Learning Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.13043 2022-09-12 cs.SD cs.CV eess.AS 62%

AudioCLIP: Extending CLIP to Image, Text and Audio

Andrey Guzhov, Federico Raue, Jörn Hees, Andreas Dengel

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、eess.AS

Comments submitted to GCPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01186 2022-09-07 cs.CL cs.AI 62%

Text-based automatic personality prediction: A bibliographic review

Ali-Reza Feizi-Derakhshi, Mohammad-Reza Feizi-Derakhshi, Majid Ramezani, Narjes Nikzad-Khasmakhi, Meysam Asgari-Chenaghlu, Taymaz Akan, Mehrdad Ranjbar-Khadivi, Elnaz Zafarni-Moattar, Zoleikha Jahanbakhsh-Naghadeh

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments This is a preprint of an article published in "Journal of Computational Social Science". The final authenticated version is available online at: https://doi.org/10.1007/s42001-022-00178-4

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14884 2022-09-01 cs.CL cs.AI 62%

GRILLBot: An Assistant for Real-World Tasks with Neural Semantic Parsing and Graph-Based Representations

Carlos Gemmell, Iain Mackie, Paul Owoicho, Federico Rossetto, Sophie Fischer, Jeffrey Dalton

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.13259 2022-08-30 cs.CL cs.AI 62%

Bayesian Neural Network Language Modeling for Speech Recognition

Boyang Xue, Shoukang Hu, Junhao Xu, Mengzhe Geng, Xunying Liu, Helen Meng

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12415 2022-08-29 eess.AS cs.CL cs.SD stat.ML 62%

MuLan: A Joint Embedding of Music Audio and Natural Language

Qingqing Huang, Aren Jansen, Joonseok Lee, Ravi Ganti, Judith Yue Li, Daniel P. W. Ellis

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments To appear in ISMIR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11671 2022-08-25 cs.SD cs.CL eess.AS 62%

Interpreting Song Lyrics with an Audio-Informed Pre-trained Language Model

Yixiao Zhang, Junyan Jiang, Gus Xia, Simon Dixon

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted to ISMIR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.02041 2022-08-23 cs.SD cs.CL eess.AS 62%

A Comparative Study of Speaker Role Identification in Air Traffic Communication Using Deep Learning Approaches

Dongyue Guo, Jianwei Zhang, Bo Yang, Yi Lin

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、eess.AS

Comments This work has been submitted to the ACM TALLIP for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.09269 2022-08-22 eess.SP cs.AI cs.LG cs.SD eess.AS 62%

Feature Selection Enhancement and Feature Space Visualization for Speech-Based Emotion Recognition

Sofia Kanwal, Sohail Asghar, Hazrat Ali

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS

Comments Accepted at PeerJ Computer Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.02058 2022-08-17 cs.SD cs.CV cs.LG eess.AS 62%

SVTS: Scalable Video-to-Speech Synthesis

Rodrigo Mira, Alexandros Haliassos, Stavros Petridis, Björn W. Schuller, Maja Pantic

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments accepted to INTERSPEECH 2022 (Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03043 2022-07-29 cs.SD cs.AI cs.LG eess.AS 62%

Sound2Synth: Interpreting Sound via FM Synthesizer Parameters Estimation

Zui Chen, Yansen Jing, Shengcheng Yuan, Yifei Xu, Jian Wu, Hang Zhao

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS

Comments 8 pages, 8 figures. v2: IJCAI2022 published, format revisions and bugfixes

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00604 2022-07-20 cs.CV cs.SD eess.AS 62%

Quantized GAN for Complex Music Generation from Dance Videos

Ye Zhu, Kyle Olszewski, Yu Wu, Panos Achlioptas, Menglei Chai, Yan Yan, Sergey Tulyakov

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、eess.AS

Comments Dataset and code at https://github.com/L-YeZhu/D2M-GAN

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.03682 2022-07-11 cs.CV cs.MM 62%

Music-driven Dance Regeneration with Controllable Key Pose Constraints

Junfu Pu, Ying Shan

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.15400 2022-07-04 eess.AS cs.AI cs.LG 62%

Learning Audio-Text Agreement for Open-vocabulary Keyword Spotting

Hyeon-Kyeong Shin, Hyewon Han, Doyeon Kim, Soo-Whan Chung, Hong-Goo Kang

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted to Interspeech 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.08039 2022-06-17 cs.SD cs.CL cs.LG eess.AS 62%

Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History

Yuto Nishimura, Yuki Saito, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments 5 pages, 3 figures, Accepted for INTERSPEECH2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.03809 2022-06-10 cs.SD cs.GR cs.MM eess.AS 62%

Music2Video: Automatic Generation of Music Video with fusion of audio and text

Yoonjeon Kim, Joel Jang, Sumin Shin

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.14769 2022-05-31 cs.CL cs.AI 62%

UPB at SemEval-2022 Task 5: Enhancing UNITER with Image Sentiment and Graph Convolutional Networks for Multimedia Automatic Misogyny Identification

Andrei Paraschiv, Mihai Dascalu, Dumitru-Clementin Cercel

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Semeval 2022, Task 5 submission 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.07205 2022-05-25 eess.AS cs.CL cs.LG cs.SD 62%

SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted by ACL 2022 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08866 2022-05-19 cs.MM cs.SD eess.AS 62%

Seeing Sounds, Hearing Shapes: a gamified study to evaluate sound-sketches

Sebastian Löbbers, György Fazekas

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM、eess.AS

Comments Accepted at International Computer Music Conference (ICMC) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.05586 2022-05-12 eess.AS cs.CV cs.LG cs.SD 62%

End-to-End Multi-Person Audio/Visual Automatic Speech Recognition

Otavio Braga, Takaki Makino, Olivier Siohan, Hank Liao

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.09737 2022-04-18 cs.CL eess.AS 62%

Consecutive Decoding for Speech-to-text Translation

Qianqian Dong, Mingxuan Wang, Hao Zhou, Shuang Xu, Bo Xu, Lei Li

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted by AAAI 2021, 11 pages, 3 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15326 2022-03-30 cs.SD cs.AI eess.AS 62%

Speech Emotion Recognition with Co-Attention based Multi-level Acoustic Information

Heqing Zou, Yuke Si, Chen Chen, Deepu Rajan, Eng Siong Chng

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments Accepted by ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15183 2022-03-30 eess.AS cs.CL cs.SD 62%

Visualizations of Complex Sequences of Family-Infant Vocalizations Using Bag-of-Audio-Words Approach Based on Wav2vec 2.0 Features

Jialu Li, Mark Hasegawa-Johnson, Nancy L. McElwain

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、eess.AS

Comments Submitted to Interspeech 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.12829 2022-03-25 cs.CV cs.MM 62%

AIMusicGuru: Music Assisted Human Pose Correction

Snehesh Shrestha, Cornelia Fermüller, Tianyu Huang, Pyone Thant Win, Adam Zukerman, Chethan M. Parameshwara, Yiannis Aloimonos

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments 10 pages, 7 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.10274 2022-03-22 eess.AS cs.AI 62%

Exploiting Cross Domain Acoustic-to-articulatory Inverted Features For Disordered Speech Recognition

Shujie Hu, Shansong Liu, Xurong Xie, Mengzhe Geng, Tianzi Wang, Shoukang Hu, Mingyu Cui, Xunying Liu, Helen Meng

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS

Comments accepted by ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.01205 2022-03-03 cs.SD cs.AI eess.AS 62%

Audio Self-supervised Learning: A Survey

Shuo Liu, Adria Mallol-Ragolta, Emilia Parada-Cabeleiro, Kun Qian, Xin Jing, Alexander Kathan, Bin Hu, Bjoern W. Schuller

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.05845 2022-03-01 eess.AS cs.AI cs.LG cs.SD 62%

Recent Progress in the CUHK Dysarthric Speech Recognition System

Shansong Liu, Mengzhe Geng, Shoukang Hu, Xurong Xie, Mingyu Cui, Jianwei Yu, Xunying Liu, Helen Meng

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.03896 2022-02-09 cs.SD cs.AI cs.LG eess.AS 62%

Speech Emotion Recognition using Self-Supervised Features

Edmilson Morais, Ron Hoory, Weizhong Zhu, Itai Gat, Matheus Damasceno, Hagai Aronowitz

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 5 pages, 4 figures, 2 tables, ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.12854 2022-01-03 cs.SD cs.MM eess.AS 62%

Audio-to-Score Alignment Using Deep Automatic Music Transcription

Federico Simonetta, Stavros Ntalampiras, Federico Avanzini

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS

Comments IEEE MMSP 2021 - ERRATUM

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08432 2021-12-17 cs.MM cs.SD eess.AS 62%

Expert and Crowd-Guided Affect Annotation and Prediction

Ramanathan Subramanian, Yan Yan, Nicu Sebe

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM、eess.AS

Comments Manuscript submitted for review to IEEE Transactions on Affective Computing

详情

展开后加载摘要…

URL PDF HTML 收藏