arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2506.16020 2025-06-23 cs.SD eess.AS 57%

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge

Zijing Zhao, Kai Wang, Hao Huang, Ying Hu, Liang He, Jichen Yang

机构 * School of Computer Science(计算机科学学院) Department of Electronic Engineering(电子工程系) School of Cyber Security(网络安全学院)

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted by Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14495 2025-06-18 cs.CV 57%

I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs

Yu Qi, Lipeng Gu, Honghua Chen, Liangliang Nan, Mingqiang Wei

机构 * School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院,南京航空航天大学) Urban Data Science Section, Delft University of Technology(都市数据科学部门,代尔夫特理工大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13990 2025-06-18 cs.AI 57%

Machine Mirages: Defining the Undefined

Hamidou Tembine

机构 * Department of Electrical Engineering and Computer Science, School of Engineering, UQTR, Quebec, Canada(电气工程与计算机科学系,工程学院,UQTR,魁北克,加拿大)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19510 2025-06-17 cs.CL 57%

Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning

Yexing Du, Youcheng Pan, Ziyang Ma, Bo Yang, Yifan Yang, Keqi Deng, Xie Chen, Yang Xiang, Ming Liu, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Pengcheng Laboratory(鹏城实验室) Shanghai Jiao Tong University(上海交通大学) University of Cambridge(剑桥大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted in ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11344 2025-06-16 cs.CL 57%

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models

Peilin Wu, Jinho D. Choi

机构 * Emory University(埃默里大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09549 2025-06-12 eess.AS cs.SD eess.SP 57%

A Study on Speech Assessment with Visual Cues

Shafique Ahmed, Ryandhimas E. Zezario, Nasir Saleem, Amir Hussain, Hsin-Min Wang, Yu Tsao

机构 * Research Center for Information Technology Innovation(信息技术创新研究中心) Institute of Information Science(信息科学研究院)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted to Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05335 2025-06-10 cs.SD eess.AS 57%

FLAM: Frame-Wise Language-Audio Modeling

Yusong Wu, Christos Tsirigotis, Ke Chen, Cheng-Zhi Anna Huang, Aaron Courville, Oriol Nieto, Prem Seetharaman, Justin Salamon

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Accepted at ICML 2025 V2: fixed small typo on eq. 15 and eq. 17

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12408 2025-06-06 cs.CL 57%

On the Robust Approximation of ASR Metrics

Abdul Waheed, Hanin Atwany, Rita Singh, Bhiksha Raj

机构 * Carnegie Mellon University(卡内基梅隆大学) MBZUAI(穆斯林人工智能研究所)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments ACL 2025 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20156 2025-06-04 cs.CV 57%

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Yi Chen, Sen Liang, Zixiang Zhou, Ziyao Huang, Yifeng Ma, Junshu Tang, Qin Lin, Yuan Zhou, Qinglin Lu

机构 * Tencent(腾讯)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01934 2025-06-03 cs.AI 57%

RoboEgo System Card: An Omnimodal Model with Native Full Duplexity

Yiqun Yao, Xiang Li, Xin Jiang, Xuezhi Fang, Naitong Yu, Aixin Sun, Yequan Wang

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院) School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01808 2025-06-03 cs.CL 57%

NAVER LABS Europe Submission to the Instruction-following Track

Beomseok Lee, Marcely Zanon Boito, Laurent Besacier, Ioan Calapodescu

机构 * NAVER LABS Europe(NAVER LABS欧洲) University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24561 2025-06-02 cs.CL 57%

Improving Language and Modality Transfer in Translation by Character-level Modeling

Ioannis Tsiamas, David Dale, Marta R. Costa-jussà

机构 * FAIR at Meta, Paris(Meta巴黎FAIR实验室) Universitat Politècnica de Catalunya, Barcelona(巴塞罗那理工大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03947 2025-05-30 cs.SD eess.AS 57%

Can Audio Reveal Music Performance Difficulty? Insights from the Piano Syllabus Dataset

Pedro Ramoneda, Minhee Lee, Dasaem Jeong, J. J. Valero-Mas, Xavier Serra

机构 * Music Technology Group, Universitat Pompeu Fabra(庞培法华大学音乐技术小组) Music & Art Learning Lab, Sogang University(ソガン大学音乐与艺术学习实验室) Pattern Recognition and Artificial Intelligence Group, University of Alicante(阿尔基兰特大学模式识别与人工智能小组)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13880 2025-05-28 eess.AS cs.SD eess.SP 57%

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

Ziqian Wang, Xianjun Xia, Xinfa Zhu, Lei Xie

机构 * Northwestern Polytechnical University(西北工业大学) Bytedance Inc.(字节跳动公司)

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments Accepted to Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19437 2025-05-27 cs.SD eess.AS 57%

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

Haoqin Sun, Jingguang Tian, Jiaming Zhou, Hui Wang, Jiabei He, Shiwan Zhao, Xiangyu Kong, Desheng Hu, Xinkang Xu, Xinhui Hu, Yong Qin

机构 * Nankai University(南开大学) TMCC, College of Computer Science(TMCC计算机学院) Hithink RoyalFlush AI Research Institute(Hithink RoyalFlush人工智能研究院) University of Exeter(埃克塞特大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18864 2025-05-27 cs.CL 57%

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework

Binhao Ma, Hanqing Guo, Zhengping Jay Luo, Rui Duan

机构 * Department of Computer Science University of Missouri-Kansas City(计算机科学系 密苏里大学-康科特分校) Department of Computer Science and Physics Rider University(计算机科学与物理系 Rider大学) Department of Electrical and Computer Engineering University of Hawai’i at Mānoa(电气与计算机工程系 夏威夷大学马诺亚分校)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16387 2025-05-23 eess.AS 57%

Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge

Ming Cheng, Fei Su, Cancan Li, Juan Liu, Ming Li

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by Interspeech2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15000 2025-05-22 cs.CL 57%

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems

Chengwei Wei, Bin Wang, Jung-jae Kim, Nancy F. Chen

机构 * Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息与通信研究 institute(I2R),A*STAR,新加坡) Centre for Frontier AI Research (CFAR), A*STAR, Singapore(前沿人工智能研究 centre(CFAR),A*STAR,新加坡)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07461 2025-05-19 cs.SD cs.AI 57%

JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata

Abhinaba Roy, Renhang Liu, Tongyu Lu, Dorien Herremans

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13933 2025-05-16 cs.CV 57%

Unsupervised Video Highlight Detection by Learning from Audio and Visual Recurrence

Zahidul Islam, Sujoy Paul, Mrigank Rochan

机构 * University of Saskatchewan(萨斯喀彻温大学) Google DeepMind(谷歌DeepMind)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Accepted to the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03244 2025-05-14 cs.SD eess.AS 57%

SonicRAG : High Fidelity Sound Effects Synthesis Based on Retrival Augmented Generation

Yu-Ren Guo, Wen-Kai Tai

机构 * Dept. of Computer Science and Information Engineering(计算机科学与信息工程系) National Taiwan University of Science and Technology(台湾科技大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16267 2025-05-14 cs.SD cs.LG eess.AS q-bio.QM 57%

A Classification Benchmark for Artificial Intelligence Detection of Laryngeal Cancer from Patient Voice

Mary Paterson, James Moor, Luisa Cutillo

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 16 pages, 6 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15118 2025-05-13 cs.CV cs.SD 57%

Improving Sound Source Localization with Joint Slot Attention on Image and Audio

Inho Kim, Youngkil Song, Jicheol Park, Won Hwa Kim, Suha Kwak

机构 * Dept. of CSE, POSTECH(计算机科学与工程系,POSTECH)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09089 2025-05-13 cs.SD cs.HC eess.AS 57%

Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter

Yi Xiao, Harshit Sharma, Victoria Tumanova, Asif Salekin

机构 * Arizona State University(亚利桑那州立大学) Syracuse University(Syracuse大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments 13 pages, 5 figures, ICCPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05755 2025-05-12 cs.SD cs.LG eess.AS 57%

CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech

Madhurananda Pahar, Fuxiang Tao, Bahman Mirheidari, Nathan Pevy, Rebecca Bright, Swapnil Gadgil, Lise Sproson, Dorota Braun, Caitlin Illingworth, Daniel Blackburn, Heidi Christensen

机构 * School of Computer Science, University of Sheffield, Sheffield, S1 4DP, UK(1 计算机科学学院,谢菲尔德大学,谢菲尔德,S1 4DP,英国) Therapy Box, London, UK(2 Therapy Box,伦敦,英国) NIHR Devices for Dignity HTC, Sheffield Teaching Hospitals NHS Foundation Trust, Sheffield, S10 2JF, UK(3 NIHR尊严设备HTC,谢菲尔德教学医院国家健康服务基金会信托,谢菲尔德,S10 2JF,英国) Sheffield Institute for Translational Neuroscience (SITraN), University of Sheffield, Sheffield, S10 2HQ, UK(4 谢菲尔德转化神经科学研究所(SITraN),谢菲尔德大学,谢菲尔德,S10 2HQ,英国)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments This paper has been accepted for publication in IEEE SSCI 2025. Copyright belongs to IEEE

Journal ref IEEE SSCI, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03033 2025-05-07 cs.AI cs.HC 57%

Evaluating the Impact of AI-Powered Audiovisual Personalization on Learner Emotion, Focus, and Learning Outcomes

George Xi Wang, Jingying Deng, Safinah Ali

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00525 2025-05-02 eess.IV cs.CV cs.LG 57%

A Methodological and Structural Review of Parkinsons Disease Detection Across Diverse Data Modalities

Abu Saleh Musa Miah, taro Suzuki, Jungpil Shin

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04644 2025-05-01 eess.AS cs.SD 57%

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Hanzhao Li, Yuke Li, Xinsheng Wang, Jingbin Hu, Qicong Xie, Shan Yang, Lei Xie

机构 * Hong Kong University of Science and Technology(香港科技大学) Tencent AI Lab(腾讯AI实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 14 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20016 2025-04-29 cs.HC cs.CY cs.MM 57%

Applying LLM-Powered Virtual Humans to Child Interviews in Child-Centered Design

Linshi Li, Hanlin Cai

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM

Comments This paper has been accepted as a Work-in-Progress (WiP) paper in the 24th annual ACM Interaction Design and Children (IDC) Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01089 2025-04-29 cs.RO cs.AI 57%

HomeEmergency -- Using Audio to Find and Respond to Emergencies in the Home

James F. Mullen, Dhruva Kumar, Xuewei Qi, Rajasimman Madhivanan, Arnie Sen, Dinesh Manocha, Richard Kim

机构 * Amazon Lab126(亚马逊实验室126) University of Maryland(马里兰大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Journal ref IEEE Robotics and Automation Letters (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏