arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2309.05248 2023-09-15 eess.AS cs.SD 57%

Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach

Tae Jin Park, Kunal Dhawan, Nithin Koluguri, Jagadeesh Balam

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments 4 pages 1 reference page, ICASSP format

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.07378 2023-09-11 eess.AS cs.LG cs.SD 57%

Dawn of the transformer era in speech emotion recognition: closing the valence gap

Johannes Wagner, Andreas Triantafyllopoulos, Hagen Wierstorf, Maximilian Schmitt, Felix Burkhardt, Florian Eyben, Björn W. Schuller

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Journal ref in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 9, pp. 10745-10759, 1 Sept. 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.08356 2023-09-07 cs.CV 57%

Leveraging TCN and Transformer for effective visual-audio fusion in continuous emotion recognition

Weiwei Zhou, Jiada Lu, Zhaolong Xiong, Weifeng Wang

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00030 2023-09-04 cs.CV 57%

Audio-Driven Dubbing for User Generated Contents via Style-Aware Semi-Parametric Synthesis

Linsen Song, Wayne Wu, Chaoyou Fu, Chen Change Loy, Ran He

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments TCSVT 2022

Journal ref IEEE Transactions on Circuits and Systems for Video Technology (Volume: 33, Issue: 3, March 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14482 2023-08-29 cs.CL 57%

An Empirical Study of Consistency Regularization for End-to-End Speech-to-Text Translation

Pengzhi Gao, Ruiqing Zhang, Zhongjun He, Hua Wu, Haifeng Wang

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14319 2023-08-29 cs.SD eess.AS 57%

Voice Conversion with Denoising Diffusion Probabilistic GAN Models

Xulong Zhang, Jianzong Wang, Ning Cheng, Jing Xiao

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by 19th International Conference on Advanced Data Mining and Applications. (ADMA 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.02053 2023-08-29 cs.CV 57%

Day2Dark: Pseudo-Supervised Activity Recognition beyond Silent Daylight

Yunhua Zhang, Hazel Doughty, Cees G. M. Snoek

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09323 2023-08-25 cs.CV 57%

Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis

Jiahe Li, Jiawei Zhang, Xiao Bai, Jun Zhou, Lin Gu

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV 2023. Project page: https://fictionarry.github.io/ER-NeRF/

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09622 2023-08-21 cs.CV 57%

Is context all you need? Scaling Neural Sign Language Translation to Large Domains of Discourse

Ozge Mercanoglu Sincan, Necati Cihan Camgoz, Richard Bowden

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03432 2023-08-08 cs.MM 57%

Cuing Without Sharing: A Federated Cued Speech Recognition Framework via Mutual Knowledge Distillation

Yuxuan Zhang, Lei Liu, Li Liu

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.02665 2023-08-08 cs.AI 57%

Let's Give a Voice to Conversational Agents in Virtual Reality

Michele Yin, Gabriel Roccabruna, Abhinav Azad, Giuseppe Riccardi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.04214 2023-08-08 eess.IV cs.CV cs.LG 57%

Intelligent Sight and Sound: A Chronic Cancer Pain Dataset

Catherine Ordun, Alexandra N. Cha, Edward Raff, Byron Gaskin, Alex Hanson, Mason Rule, Sanjay Purushotham, James L. Gulley

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Published as conference paper at the 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks

Journal ref 2021, Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10741 2023-08-04 cs.CV cs.LG 57%

Computer Vision Estimation of Emotion Reaction Intensity in the Wild

Yang Qian, Ali Kargarandehkordi, Onur Cezmi Mutlu, Saimourya Surabhi, Mohammadmahdi Honarmand, Dennis Paul Wall, Peter Washington

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14739 2023-07-28 eess.AS eess.SP 57%

Audio Inputs for Active Speaker Detection and Localization via Microphone Array

Davide Berghi, Philip J. B. Jackson

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.10008 2023-07-20 cs.CV 57%

MODA: Mapping-Once Audio-driven Portrait Animation with Dual Attentions

Yunfei Liu, Lijian Lin, Fei Yu, Changyin Zhou, Yu Li

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09006 2023-07-19 cs.SD cs.LG eess.AS 57%

OxfordVGG Submission to the EGO4D AV Transcription Challenge

Jaesung Huh, Max Bain, Andrew Zisserman

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06385 2023-07-12 cs.CL 57%

TencentPretrain: A Scalable and Flexible Toolkit for Pre-training Models of Different Modalities

Zhe Zhao, Yudong Li, Cheng Hou, Jing Zhao, Rong Tian, Weijie Liu, Yiren Chen, Ningyuan Sun, Haoyan Liu, Weiquan Mao, Han Guo, Weigang Guo, Taiqiang Wu, Tao Zhu, Wenhang Shi, Chen Chen, Shan Huang, Sihong Chen, Liqun Liu, Feifei Li, Xiaoshuai Chen, Xingwu Sun, Zhanhui Kang, Xiaoyong Du, Linlin Shen, Kimmo Yan

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.16083 2023-06-29 cs.SD eess.AS 57%

UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data

Heeseung Kim, Sungwon Kim, Jiheum Yeom, Sungroh Yoon

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments INTERSPEECH 2023, Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.13662 2023-06-27 cs.SD eess.AS 57%

InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Dongchao Yang, Songxiang Liu, Rongjie Huang, Chao Weng, Helen Meng

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments Submit to TASLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.10772 2023-06-21 cs.SD eess.AS 57%

Learning an Interpretable End-to-End Network for Real-Time Acoustic Beamforming

Hao Liang, Guanxing Zhou, Xiaotong Tu, Andreas Jakobsson, Xinghao Ding, Yue Huang

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.10053 2023-06-21 cs.IR cs.AI econ.GN q-fin.EC 57%

NFTs to MARS: Multi-Attention Recommender System for NFTs

Seonmi Kim, Youngbin Lee, Yejin Kim, Joohwan Hong, Yongjae Lee

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.07744 2023-06-14 cs.SD cs.LG eess.AS 57%

Contrastive Learning-Based Audio to Lyrics Alignment for Multiple Languages

Simon Durand, Daniel Stoller, Sebastian Ewert

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments 5 pages, accepted at the International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2023

Journal ref ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 2023, pp. 1-5

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19522 2023-06-02 cs.SD eess.AS 57%

PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions

Guanghou Liu, Yongmao Zhang, Yi Lei, Yunlin Chen, Rui Wang, Zhifei Li, Lei Xie

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.00359 2023-05-31 eess.AS 57%

A Review of Deep Learning Techniques for Speech Processing

Ambuj Mehrish, Navonil Majumder, Rishabh Bhardwaj, Rada Mihalcea, Soujanya Poria

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17137 2023-05-30 cs.AI cs.LG 57%

Integrating Generative Artificial Intelligence in Intelligent Vehicle Systems

Lukas Stappen, Jeremy Dillmann, Serena Striegel, Hans-Jörg Vögel, Nicolas Flores-Herr, Björn W. Schuller

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.05534 2023-05-10 cs.CV cs.LG 57%

Integrating Holistic and Local Information to Estimate Emotional Reaction Intensity

Yini Fang, Liang Wu, Frederic Jumelle, Bertram Shi

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments This paper will be published in CVPRW 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.01241 2023-05-09 cs.HC cs.GR cs.LG cs.SD eess.AS 57%

AQ-GT: a Temporally Aligned and Quantized GRU-Transformer for Co-Speech Gesture Synthesis

Hendric Voß, Stefan Kopp

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.00537 2023-05-02 cs.MM cs.CY cs.LG 57%

Interpretability of Machine Learning: Recent Advances and Future Prospects

Lei Gao, Ling Guan

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM

Comments IEEE Multimedia (Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.10871 2023-05-02 cs.LG cs.CL cs.SI 57%

Qualitative Analysis of a Graph Transformer Approach to Addressing Hate Speech: Adapting to Dynamically Changing Content

Liam Hebert, Hong Yi Chen, Robin Cohen, Lukasz Golab

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

Comments Accepted at AAAI 2023 AI for Social Good

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.06786 2023-04-28 eess.AS cs.SD 57%

The future of hearing aid technology

Volker Hohmann

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏