arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2306.05374 2023-10-19 physics.med-ph cs.SD eess.AS eess.IV 57%

Towards Ultrasound Tongue Image prediction from EEG during speech production

Tamás Gábor Csapó, Frigyes Viktor Arthur, Péter Nagy, Ádám Boncz

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments accepted at Interspeech 2023

Journal ref Proceedings of Interspeech 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11295 2023-10-18 cs.CV cs.CG 57%

CorrTalk: Correlation Between Hierarchical Speech and Facial Activity Variances for 3D Animation

Zhaojie Chu, Kailing Guo, Xiaofen Xing, Yilin Lan, Bolun Cai, Xiangmin Xu

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10179 2023-10-17 eess.AS cs.SD 57%

Advancing Audio Emotion and Intent Recognition with Large Pre-Trained Models and Bayesian Inference

Dejan Porjazovski, Yaroslav Getman, Tamás Grósz, Mikko Kurimo

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted at ACMM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08749 2023-10-11 cs.CL cs.LG 57%

Utilizing Longitudinal Chest X-Rays and Reports to Pre-Fill Radiology Reports

Qingqing Zhu, Tejas Sudharshan Mathai, Pritam Mukherjee, Yifan Peng, Ronald M. Summers, Zhiyong Lu

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05304 2023-10-10 cs.CV 57%

GestSync: Determining who is speaking without a talking head

Sindhu B Hegde, Andrew Zisserman

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Accepted in BMVC 2023, 10 pages paper, 7 pages supplementary, 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11096 2023-10-09 cs.SD cs.LG eess.AS 57%

Robust One-Shot Singing Voice Conversion

Naoya Takahashi, Mayank Kumar Singh, Yuki Mitsufuji

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02419 2023-10-03 cs.CV 57%

TM2D: Bimodality Driven 3D Dance Generation via Music-Text Integration

Kehong Gong, Dongze Lian, Heng Chang, Chuan Guo, Zihang Jiang, Xinxin Zuo, Michael Bi Mi, Xinchao Wang

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12111 2023-09-22 cs.SD cs.IR cs.LG eess.AS 57%

Passage Summarization with Recurrent Models for Audio-Sheet Music Retrieval

Luis Carvalho, Gerhard Widmer

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments In Proceedings of the 24th Conference of the International Society for Music Information Retrieval (ISMIR 2023), Milan, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05248 2023-09-15 eess.AS cs.SD 57%

Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach

Tae Jin Park, Kunal Dhawan, Nithin Koluguri, Jagadeesh Balam

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments 4 pages 1 reference page, ICASSP format

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.07378 2023-09-11 eess.AS cs.LG cs.SD 57%

Dawn of the transformer era in speech emotion recognition: closing the valence gap

Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Felix Burkhardt, Florian Eyben, Björn W. Schuller

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Journal ref in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 9, pp. 10745-10759, 1 Sept. 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08356 2023-09-07 cs.CV 57%

Leveraging TCN and Transformer for effective visual-audio fusion in continuous emotion recognition

Weiwei Zhou, Jiada Lu, Zhaolong Xiong, Weifeng Wang

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00030 2023-09-04 cs.CV 57%

Audio-Driven Dubbing for User Generated Contents via Style-Aware Semi-Parametric Synthesis

Linsen Song, Wayne Wu, Chaoyou Fu, Chen Change Loy, Ran He

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments TCSVT 2022

Journal ref IEEE Transactions on Circuits and Systems for Video Technology (Volume: 33, Issue: 3, March 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14482 2023-08-29 cs.CL 57%

An Empirical Study of Consistency Regularization for End-to-End Speech-to-Text Translation

Pengzhi Gao, Ruiqing Zhang, Zhongjun He, Hua Wu, Haifeng Wang

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14319 2023-08-29 cs.SD eess.AS 57%

Voice Conversion with Denoising Diffusion Probabilistic GAN Models

Xulong Zhang, Jianzong Wang, Ning Cheng, Jing Xiao

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by 19th International Conference on Advanced Data Mining and Applications. (ADMA 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.02053 2023-08-29 cs.CV 57%

Day2Dark: Pseudo-Supervised Activity Recognition beyond Silent Daylight

Yunhua Zhang, Hazel Doughty, Cees G. M. Snoek

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09323 2023-08-25 cs.CV 57%

Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis

Jiahe Li, Jiawei Zhang, Xiao Bai, Jun Zhou, Lin Gu

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2023. Project page: https://fictionarry.github.io/ER-NeRF/

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09622 2023-08-21 cs.CV 57%

Is context all you need? Scaling Neural Sign Language Translation to Large Domains of Discourse

Ozge Mercanoglu Sincan, Necati Cihan Camgoz, Richard Bowden

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03432 2023-08-08 cs.MM 57%

Cuing Without Sharing: A Federated Cued Speech Recognition Framework via Mutual Knowledge Distillation

Yuxuan Zhang, Lei Liu, Li Liu

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02665 2023-08-08 cs.AI 57%

Let's Give a Voice to Conversational Agents in Virtual Reality

Michele Yin, Gabriel Roccabruna, Abhinav Azad, Giuseppe Riccardi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.04214 2023-08-08 eess.IV cs.CV cs.LG 57%

Intelligent Sight and Sound: A Chronic Cancer Pain Dataset

Catherine Ordun, Alexandra N. Cha, Edward Raff, Byron Gaskin, Alex Hanson, Mason Rule, Sanjay Purushotham, James L. Gulley

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Published as conference paper at the 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks

Journal ref 2021, Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10741 2023-08-04 cs.CV cs.LG 57%

Computer Vision Estimation of Emotion Reaction Intensity in the Wild

Yang Qian, Ali Kargarandehkordi, Onur Cezmi Mutlu, Saimourya Surabhi, Mohammadmahdi Honarmand, Dennis Paul Wall, Peter Washington

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14739 2023-07-28 eess.AS eess.SP 57%

Audio Inputs for Active Speaker Detection and Localization via Microphone Array

Davide Berghi, Philip J. B. Jackson

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10008 2023-07-20 cs.CV 57%

MODA: Mapping-Once Audio-driven Portrait Animation with Dual Attentions

Yunfei Liu, Lijian Lin, Fei Yu, Changyin Zhou, Yu Li

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09006 2023-07-19 cs.SD cs.LG eess.AS 57%

OxfordVGG Submission to the EGO4D AV Transcription Challenge

Jaesung Huh, Max Bain, Andrew Zisserman

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06385 2023-07-12 cs.CL 57%

TencentPretrain: A Scalable and Flexible Toolkit for Pre-training Models of Different Modalities

Zhe Zhao, Yudong Li, Cheng Hou, Jing Zhao, Rong Tian, Weijie Liu, Yiren Chen, Ningyuan Sun, Haoyan Liu, Weiquan Mao, Han Guo, Weigang Guo, Taiqiang Wu, Tao Zhu, Wenhang Shi, Chen Chen, Shan Huang, Sihong Chen, Liqun Liu, Feifei Li, Xiaoshuai Chen, Xingwu Sun, Zhanhui Kang, Xiaoyong Du, Linlin Shen, Kimmo Yan

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16083 2023-06-29 cs.SD eess.AS 57%

UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data

Heeseung Kim, Sungwon Kim, Jiheum Yeom, Sungroh Yoon

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments INTERSPEECH 2023, Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13662 2023-06-27 cs.SD eess.AS 57%

InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Dongchao Yang, Songxiang Liu, Rongjie Huang, Chao Weng, Helen Meng

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments Submit to TASLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.10772 2023-06-21 cs.SD eess.AS 57%

Learning an Interpretable End-to-End Network for Real-Time Acoustic Beamforming

Hao Liang, Guanxing Zhou, Xiaotong Tu, Andreas Jakobsson, Xinghao Ding, Yue Huang

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.10053 2023-06-21 cs.IR cs.AI econ.GN q-fin.EC 57%

NFTs to MARS: Multi-Attention Recommender System for NFTs

Seonmi Kim, Youngbin Lee, Yejin Kim, Joohwan Hong, Yongjae Lee

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07744 2023-06-14 cs.SD cs.LG eess.AS 57%

Contrastive Learning-Based Audio to Lyrics Alignment for Multiple Languages

Simon Durand, Daniel Stoller, Sebastian Ewert

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments 5 pages, accepted at the International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2023

Journal ref ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 2023, pp. 1-5

详情

展开后加载摘要…

URL PDF HTML 收藏