arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2211.00924 2022-11-04 cs.CV cs.AI eess.IV 62%

SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory

Se Jin Park, Minsu Kim, Joanna Hong, Jeongsoo Choi, Yong Man Ro

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.AI

Comments Accepted at AAAI 2022 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00549 2022-11-02 cs.CV cs.LG cs.SD eess.AS 62%

No-audio speaking status detection in crowded settings via visual pose-based filtering and wearable acceleration

Jose Vargas-Quiros, Laura Cabrera-Quiros, Hayley Hung

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13084 2022-11-01 cs.CV cs.SD eess.AS 62%

Visual Speech Recognition for Multiple Languages in the Wild

Pingchuan Ma, Stavros Petridis, Maja Pantic

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments Published in Nature Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.15828 2022-10-31 cs.SD cs.MM eess.AS 62%

On the Role of Visual Context in Enriching Music Representations

Kleanthis Avramidis, Shanti Stewart, Shrikanth Narayanan

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS

Comments 5 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12857 2022-10-25 cs.CL cs.SD eess.AS 62%

Bootstrapping meaning through listening: Unsupervised learning of spoken sentence embeddings

Jian Zhu, Zuoyu Tian, Yadong Liu, Cong Zhang, Chia-wen Lo

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Findings of EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.06423 2022-10-20 cs.LG cs.CL cs.CV 62%

Foundation Transformers

Hongyu Wang, Shuming Ma, Shaohan Huang, Li Dong, Wenhui Wang, Zhiliang Peng, Yu Wu, Payal Bajaj, Saksham Singhal, Alon Benhaim, Barun Patra, Zhun Liu, Vishrav Chaudhary, Xia Song, Furu Wei

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.05766 2022-10-13 cs.CV cs.LG cs.MM 62%

Match Cutting: Finding Cuts with Smooth Visual Transitions

Boris Chen, Amir Ziai, Rebecca Tucker, Yuchen Xie

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03730 2022-10-10 cs.CL eess.AS 62%

SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training

Ziqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu, Lirong Dai, Jinyu Li, Furu Wei

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments 14 pages, accepted by EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.04748 2022-09-13 cs.SD cs.AI eess.AS 62%

LipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction and Lip Reading

Leyuan Qu, Cornelius Weber, Stefan Wermter

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS

Comments ACCEPTED IN IEEE Transactions on Neural Networks and Learning Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.13043 2022-09-12 cs.SD cs.CV eess.AS 62%

AudioCLIP: Extending CLIP to Image, Text and Audio

Andrey Guzhov, Federico Raue, Jörn Hees, Andreas Dengel

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、eess.AS

Comments submitted to GCPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.01186 2022-09-07 cs.CL cs.AI 62%

Text-based automatic personality prediction: A bibliographic review

Ali-Reza Feizi-Derakhshi, Mohammad-Reza Feizi-Derakhshi, Majid Ramezani, Narjes Nikzad-Khasmakhi, Meysam Asgari-Chenaghlu, Taymaz Akan, Mehrdad Ranjbar-Khadivi, Elnaz Zafarni-Moattar, Zoleikha Jahanbakhsh-Naghadeh

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments This is a preprint of an article published in "Journal of Computational Social Science". The final authenticated version is available online at: https://doi.org/10.1007/s42001-022-00178-4

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14884 2022-09-01 cs.CL cs.AI 62%

GRILLBot: An Assistant for Real-World Tasks with Neural Semantic Parsing and Graph-Based Representations

Carlos Gemmell, Iain Mackie, Paul Owoicho, Federico Rossetto, Sophie Fischer, Jeffrey Dalton

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.13259 2022-08-30 cs.CL cs.AI 62%

Bayesian Neural Network Language Modeling for Speech Recognition

Boyang Xue, Shoukang Hu, Junhao Xu, Mengzhe Geng, Xunying Liu, Helen Meng

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12415 2022-08-29 eess.AS cs.CL cs.SD stat.ML 62%

MuLan: A Joint Embedding of Music Audio and Natural Language

Qingqing Huang, Aren Jansen, Joonseok Lee, Ravi Ganti, Judith Yue Li, Daniel P. W. Ellis

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments To appear in ISMIR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11671 2022-08-25 cs.SD cs.CL eess.AS 62%

Interpreting Song Lyrics with an Audio-Informed Pre-trained Language Model

Yixiao Zhang, Junyan Jiang, Gus Xia, Simon Dixon

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted to ISMIR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.02041 2022-08-23 cs.SD cs.CL eess.AS 62%

A Comparative Study of Speaker Role Identification in Air Traffic Communication Using Deep Learning Approaches

Dongyue Guo, Jianwei Zhang, Bo Yang, Yi Lin

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、eess.AS

Comments This work has been submitted to the ACM TALLIP for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.09269 2022-08-22 eess.SP cs.AI cs.LG cs.SD eess.AS 62%

Feature Selection Enhancement and Feature Space Visualization for Speech-Based Emotion Recognition

Sofia Kanwal, Sohail Asghar, Hazrat Ali

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS

Comments Accepted at PeerJ Computer Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.02058 2022-08-17 cs.SD cs.CV cs.LG eess.AS 62%

SVTS: Scalable Video-to-Speech Synthesis

Rodrigo Mira, Alexandros Haliassos, Stavros Petridis, Björn W. Schuller, Maja Pantic

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

Comments accepted to INTERSPEECH 2022 (Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03043 2022-07-29 cs.SD cs.AI cs.LG eess.AS 62%

Sound2Synth: Interpreting Sound via FM Synthesizer Parameters Estimation

Zui Chen, Yansen Jing, Shengcheng Yuan, Yifei Xu, Jian Wu, Hang Zhao

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS

Comments 8 pages, 8 figures. v2: IJCAI2022 published, format revisions and bugfixes

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00604 2022-07-20 cs.CV cs.SD eess.AS 62%

Quantized GAN for Complex Music Generation from Dance Videos

Ye Zhu, Kyle Olszewski, Yu Wu, Panos Achlioptas, Menglei Chai, Yan Yan, Sergey Tulyakov

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、eess.AS

Comments Dataset and code at https://github.com/L-YeZhu/D2M-GAN

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.03682 2022-07-11 cs.CV cs.MM 62%

Music-driven Dance Regeneration with Controllable Key Pose Constraints

Junfu Pu, Ying Shan

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.15400 2022-07-04 eess.AS cs.AI cs.LG 62%

Learning Audio-Text Agreement for Open-vocabulary Keyword Spotting

Hyeon-Kyeong Shin, Hyewon Han, Doyeon Kim, Soo-Whan Chung, Hong-Goo Kang

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted to Interspeech 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.08039 2022-06-17 cs.SD cs.CL cs.LG eess.AS 62%

Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History

Yuto Nishimura, Yuki Saito, Shinnosuke Takamichi, Kentaro Tachibana, Hiroshi Saruwatari

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments 5 pages, 3 figures, Accepted for INTERSPEECH2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.03809 2022-06-10 cs.SD cs.GR cs.MM eess.AS 62%

Music2Video: Automatic Generation of Music Video with fusion of audio and text

Yoonjeon Kim, Joel Jang, Sumin Shin

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.14769 2022-05-31 cs.CL cs.AI 62%

UPB at SemEval-2022 Task 5: Enhancing UNITER with Image Sentiment and Graph Convolutional Networks for Multimedia Automatic Misogyny Identification

Andrei Paraschiv, Mihai Dascalu, Dumitru-Clementin Cercel

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments Semeval 2022, Task 5 submission 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.07205 2022-05-25 eess.AS cs.CL cs.LG cs.SD 62%

SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted by ACL 2022 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08866 2022-05-19 cs.MM cs.SD eess.AS 62%

Seeing Sounds, Hearing Shapes: a gamified study to evaluate sound-sketches

Sebastian Löbbers, György Fazekas

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM、eess.AS

Comments Accepted at International Computer Music Conference (ICMC) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.05586 2022-05-12 eess.AS cs.CV cs.LG cs.SD 62%

End-to-End Multi-Person Audio/Visual Automatic Speech Recognition

Otavio Braga, Takaki Makino, Olivier Siohan, Hank Liao

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.09737 2022-04-18 cs.CL eess.AS 62%

Consecutive Decoding for Speech-to-text Translation

Qianqian Dong, Mingxuan Wang, Hao Zhou, Shuang Xu, Bo Xu, Lei Li

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted by AAAI 2021, 11 pages, 3 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15326 2022-03-30 cs.SD cs.AI eess.AS 62%

Speech Emotion Recognition with Co-Attention based Multi-level Acoustic Information

Heqing Zou, Yuke Si, Chen Chen, Deepu Rajan, Eng Siong Chng

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments Accepted by ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏