arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46352 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4587 篇

2606.17281 2026-06-17 cs.CL cs.SD eess.AS 新提交 76%

Are you speaking my languages? On spoken language adherence in multimodal LLMs

你在说我的语言吗?多模态大语言模型中的口语遵循问题

Hyungwon Kim, Kandarp Joshi, Lillian Zhou, Pavel Golik, Petar Aleksic

机构 * Google DeepMind(谷歌DeepMind)

专题命中 音频语音多模态 :multimodal(title);分类 cs.CL、eess.AS

AI总结 针对多模态大语言模型在自动语音识别中输出语言识别错误的问题,提出软提示方法、监督微调和思维链推理三种缓解策略,并引入新指标量化语言违背,比较各方法在减少违规和保持ASR性能上的效果。

Comments 7 pages, 3 tables in the main body

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19176 2026-03-20 cs.SD cs.CV eess.AS 76%

Few-shot Acoustic Synthesis with Multimodal Flow Matching

少样本声学合成与多模态流匹配

Amandine Brunetto

机构 * Center for Robotics, Mines Paris - PSL University(机器人中心,巴黎 Mines - PSL 大学)

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、eess.AS

AI总结 本文提出FLAC方法,通过流匹配生成少样本声学合成,结合空间、几何和声学线索生成新的场景房间脉冲响应,优于现有八样本基线。

Comments To appear at CVPR 2026. 23 pages, 16 figures. Project Page: https://amandinebtto.github.io/FLAC/

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04561 2025-09-24 cs.CL cs.CV 76%

OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis

Run Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu, Xiong Liu, Min Yang, Yongbin Li, Longze Chen, Jiaming Li, Lei Zhang, Xiaobo Xia, Hamid Alinejad-Rokny, Fei Huang

机构 * Shenzhen Key Laboratory for High Performance Data Mining(深圳高性能数据挖掘重点实验室) Shenzhen Institute of Advanced Technology(深圳先进技术研究院) Chinese Academy of Sciences(中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Tongyi Laboratory(通义实验室) University of New South Wales(新南威尔士大学) National University of Singapore(新加坡国立大学) University of Science and Technology of China(中国科学技术大学) MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition(脑启发智能感知与认知重点实验室)

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17523 2025-09-23 cs.CL eess.AS 76%

Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models

María Andrea Cruz Blandón, Zakaria Aldeneh, Jie Chi, Maureen de Seyssel

机构 * Tampere University Apple(塔尔库大学苹果)

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CL、eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00033 2025-09-19 cs.CV cs.AI 76%

Deep Learning-Driven Multimodal Detection and Movement Analysis of Objects in Culinary

Tahoshin Alam Ishat, Mohammad Abdul Qayum

机构 * Electrical and Computer Engineering(电气与计算机工程) North South University(北南大学) Computer Engineering North South University Dhaka, Bangladesh(北南大学计算机工程系)

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.AI

Comments 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15810 2025-08-25 cs.CL cs.AI cs.LG 76%

Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models

Nouar AlDahoul, Yasir Zaki

机构 * Computer Science Department, New York University Abu Dhabi(纽约大学阿布扎克分校计算机科学系)

专题命中 音频语音多模态 :multi-modal(title);分类 cs.CL、cs.AI

Comments 26 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23406 2025-05-30 cs.CV cs.AI cs.LG 76%

Video Editing for Audio-Visual Dubbing

Binyamin Manela, Sharon Gannot, Ethan Fetyaya

机构 * Faculty of Engineering Bar-Ilan University(巴伊兰大学工程学院)

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00175 2025-05-30 cs.CV cs.LG cs.SD eess.AS eess.IV 76%

Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning

Stefan Smeu, Dragos-Alexandru Boldisor, Dan Oneata, Elisabeta Oneata

机构 * Bitdefender Politehnica Bucharest(巴尔干理工大学)

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、eess.AS

Comments Accepted as a highlight paper at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16691 2025-05-26 cs.SD cs.AI eess.AS 76%

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

Advait Joglekar, Divyanshu Singh, Rooshil Rohit Bhatia, S. Umesh

机构 * SPRING Lab, Indian Institute of Technology Madras(SPRING实验室,印度理工学院马德拉斯学院)

专题命中 音频语音多模态 :any-to-any(title);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13079 2025-05-20 eess.AS cs.AI 76%

Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR

Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai

专题命中 音频语音多模态 :cross-modal(title);分类 cs.AI、eess.AS

Comments To appear in Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16874 2025-04-29 cs.AI eess.AS 76%

A Multi-modal Approach to Dysarthria Detection and Severity Assessment Using Speech and Text Information

Anuprabha M, Krishna Gurugubelli, V Kesavaraj, Anil Kumar Vuppala

机构 * Speech Processing Laboratory, LTRC IIIT Hyderabad, India(国际信息科技研究所-海得拉巴语音处理实验室) Samsung Research & Development Institute-Bengaluru, India(三星研发研究所-班加罗尔)

专题命中 音频语音多模态 :multi-modal(title);分类 cs.AI、eess.AS

Comments Submitted to ICASSP 2025

Journal ref ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 2025, pp. 1-5

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17572 2024-12-24 cs.CV cs.AI 76%

Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor

Yeonju Kim, Se Jin Park, Yong Man Ro

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01620 2024-11-12 cs.SD cs.AI cs.CY eess.AS 76%

Voice EHR: Introducing Multimodal Audio Data for Health

James Anibal, Hannah Huth, Ming Li, Lindsey Hazen, Veronica Daoud, Dominique Ebedes, Yen Minh Lam, Hang Nguyen, Phuc Hong, Michael Kleinman, Shelley Ost, Christopher Jackson, Laura Sprabery, Cheran Elangovan, Balaji Krishnaiah, Lee Akst, Ioan Lina, Iqbal Elyazar, Lenny Ekwati, Stefan Jansen, Richard Nduwayezu, Charisse Garcia, Jeffrey Plum, Jacqueline Brenner, Miranda Song, Emily Ricotta, David Clifton, C. Louise Thwaites, Yael Bensoussan, Bradford Wood

专题命中 音频语音多模态 :multimodal(title);分类 cs.AI、eess.AS

Comments 21 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09928 2024-10-15 cs.SD cs.AI eess.AS 76%

M2M-Gen: A Multimodal Framework for Automated Background Music Generation in Japanese Manga Using Large Language Models

Megha Sharma, Muhammad Taimoor Haseeb, Gus Xia, Yoshimasa Tsuruoka

专题命中 音频语音多模态 :multimodal(title);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.05007 2024-09-10 cs.SD cs.AI eess.AS 76%

Audio-Guided Fusion Techniques for Multimodal Emotion Analysis

Pujin Shi, Fei Gao

专题命中 音频语音多模态 :multimodal(title);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10383 2024-08-21 cs.SD cs.AI eess.AS 76%

BrewCLIP: A Bifurcated Representation Learning Framework for Audio-Visual Retrieval

Zhenyu Lu, Lakshay Sethi

专题命中 音频语音多模态 :audio-visual(title);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12842 2024-07-19 cs.CL cs.AI 76%

MS2SL: Multimodal Spoken Data-Driven Continuous Sign Language Production

Jian Ma, Wenguan Wang, Yi Yang, Feng Zheng

专题命中 音频语音多模态 :multimodal(title);分类 cs.CL、cs.AI

Comments Accepted to ACL 2024 Findings; Project Page: https://hechang25.github.io/MS2SL

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05746 2024-07-09 cs.AI cs.SD eess.AS 76%

MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition

Jarod Duret, Mickael Rouvier, Yannick Estève

专题命中 音频语音多模态 :multimodal(title);分类 cs.AI、eess.AS

Journal ref Odyssey 2024, Jun 2024, Quebec, France

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.08317 2024-05-15 cs.CL cs.SD eess.AS 76%

SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models

Raghuveer Peri, Sai Muralidhar Jayanthi, Srikanth Ronanki, Anshu Bhatia, Karel Mundnich, Saket Dingliwal, Nilaksh Das, Zejiang Hou, Goeric Huybrechts, Srikanth Vishnubhotla, Daniel Garcia-Romero, Sundararajan Srinivasan, Kyu J Han, Katrin Kirchhoff

专题命中 音频语音多模态 :multimodal(title);分类 cs.CL、eess.AS

Comments 9+6 pages, Submitted to ACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11954 2024-02-20 cs.SD cs.MM eess.AS 76%

Multimodal Emotion Recognition from Raw Audio with Sinc-convolution

Xiaohui Zhang, Wenjie Fu, Mangui Liang

专题命中 音频语音多模态 :multimodal(title);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10653 2024-01-22 cs.CL cs.LG cs.SD eess.AS eess.SP 76%

Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection

Atanu Mandal, Gargi Roy, Amit Barman, Indranil Dutta, Sudip Kumar Naskar

专题命中 音频语音多模态 :multimodal(title);分类 cs.CL、eess.AS

Comments Accepted in 20th International Conference on Natural Language Processing (ICON)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.02392 2024-01-18 eess.IV cs.CV cs.MM 76%

Audio-Visual Quality Assessment for User Generated Content: Database and Method

Yuqin Cao, Xiongkuo Min, Wei Sun, Xiaoping Zhang, Guangtao Zhai

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08738 2023-12-22 cs.CV cs.MM 76%

AV-MaskEnhancer: Enhancing Video Representations through Audio-Visual Masked Autoencoder

Xingjian Diao, Ming Cheng, Shitong Cheng

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.MM

Comments 2023 IEEE 35th International Conference on Tools with Artificial Intelligence (ICTAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14091 2023-11-27 cs.HC cs.AI cs.CY cs.MM 76%

PortfolioMentor: Multimodal Generative AI Companion for Learning and Crafting Interactive Digital Art Portfolios

Tao Long, Weirui Peng

专题命中 音频语音多模态 :multimodal(title);分类 cs.AI、cs.MM

Comments 3 pages, 1 figure, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06737 2023-11-14 cs.CL cs.AI 76%

Detecting and Correcting Hate Speech in Multimodal Memes with Large Visual Language Model

Minh-Hao Van, Xintao Wu

专题命中 音频语音多模态 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05071 2023-11-10 cs.LG cs.CV cs.SD eess.AS 76%

On the Behavior of Audio-Visual Fusion Architectures in Identity Verification Tasks

Daniel Claborne, Eric Slyman, Karl Pazdernik

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11699 2023-10-19 cs.CL cs.CV 76%

MISAR: A Multimodal Instructional System with Augmented Reality

Jing Bi, Nguyen Manh Nguyen, Ali Vosoughi, Chenliang Xu

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.CL

Comments Accepted at ICCV 2023 - AV4D, 6 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06702 2023-10-11 cs.CL cs.LG cs.SD eess.AS 76%

Temporally Aligning Long Audio Interviews with Questions: A Case Study in Multimodal Data Integration

Piyush Singh Pasi, Karthikeya Battepati, Preethi Jyothi, Ganesh Ramakrishnan, Tanmay Mahapatra, Manoj Singh

专题命中 音频语音多模态 :multimodal(title);分类 cs.CL、eess.AS

Comments Work Accepted in IJCAI-23- AI and Social Good Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.13354 2023-09-26 cs.CL cs.AI cs.LG 76%

Lexical Squad@Multimodal Hate Speech Event Detection 2023: Multimodal Hate Speech Detection using Fused Ensemble Approach

Mohammad Kashif, Mohammad Zohair, Saquib Ali

专题命中 音频语音多模态 :multimodal(title);分类 cs.CL、cs.AI

Comments 8 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.08295 2023-09-18 eess.AS cs.CV cs.LG cs.SD 76%

A Real-Time Active Speaker Detection System Integrating an Audio-Visual Signal with a Spatial Querying Mechanism

Ilya Gurvich, Ido Leichter, Dharmendar Reddy Palle, Yossi Asher, Alon Vinnikov, Igor Abramovski, Vishak Gopal, Ross Cutler, Eyal Krupka

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏