arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2505.20156 2025-06-04 cs.CV 57%

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Yi Chen, Sen Liang, Zixiang Zhou, Ziyao Huang, Yifeng Ma, Junshu Tang, Qin Lin, Yuan Zhou, Qinglin Lu

机构 * Tencent(腾讯)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01934 2025-06-03 cs.AI 57%

RoboEgo System Card: An Omnimodal Model with Native Full Duplexity

Yiqun Yao, Xiang Li, Xin Jiang, Xuezhi Fang, Naitong Yu, Aixin Sun, Yequan Wang

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院) School of Computer Science and Engineering(计算机科学与工程学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01808 2025-06-03 cs.CL 57%

NAVER LABS Europe Submission to the Instruction-following Track

Beomseok Lee, Marcely Zanon Boito, Laurent Besacier, Ioan Calapodescu

机构 * NAVER LABS Europe(NAVER LABS欧洲) University of Trento(特伦托大学) Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24561 2025-06-02 cs.CL 57%

Improving Language and Modality Transfer in Translation by Character-level Modeling

Ioannis Tsiamas, David Dale, Marta R. Costa-jussà

机构 * FAIR at Meta, Paris(Meta巴黎FAIR实验室) Universitat Politècnica de Catalunya, Barcelona(巴塞罗那理工大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03947 2025-05-30 cs.SD eess.AS 57%

Can Audio Reveal Music Performance Difficulty? Insights from the Piano Syllabus Dataset

Pedro Ramoneda, Minhee Lee, Dasaem Jeong, J. J. Valero-Mas, Xavier Serra

机构 * Music Technology Group, Universitat Pompeu Fabra(庞培法华大学音乐技术小组) Music & Art Learning Lab, Sogang University(ソガン大学音乐与艺术学习实验室) Pattern Recognition and Artificial Intelligence Group, University of Alicante(阿尔基兰特大学模式识别与人工智能小组)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13880 2025-05-28 eess.AS cs.SD eess.SP 57%

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

Ziqian Wang, Xianjun Xia, Xinfa Zhu, Lei Xie

机构 * Northwestern Polytechnical University(西北工业大学) Bytedance Inc.(字节跳动公司)

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments Accepted to Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19437 2025-05-27 cs.SD eess.AS 57%

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

Haoqin Sun, Jingguang Tian, Jiaming Zhou, Hui Wang, Jiabei He, Shiwan Zhao, Xiangyu Kong, Desheng Hu, Xinkang Xu, Xinhui Hu, Yong Qin

机构 * Nankai University(南开大学) TMCC, College of Computer Science(TMCC计算机学院) Hithink RoyalFlush AI Research Institute(Hithink RoyalFlush人工智能研究院) University of Exeter(埃克塞特大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18864 2025-05-27 cs.CL 57%

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework

Binhao Ma, Hanqing Guo, Zhengping Jay Luo, Rui Duan

机构 * Department of Computer Science University of Missouri-Kansas City(计算机科学系 密苏里大学-康科特分校) Department of Computer Science and Physics Rider University(计算机科学与物理系 Rider大学) Department of Electrical and Computer Engineering University of Hawai’i at Mānoa(电气与计算机工程系 夏威夷大学马诺亚分校)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16387 2025-05-23 eess.AS 57%

Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge

Ming Cheng, Fei Su, Cancan Li, Juan Liu, Ming Li

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by Interspeech2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15000 2025-05-22 cs.CL 57%

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems

Chengwei Wei, Bin Wang, Jung-jae Kim, Nancy F. Chen

机构 * Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息与通信研究 institute(I2R),A*STAR,新加坡) Centre for Frontier AI Research (CFAR), A*STAR, Singapore(前沿人工智能研究 centre(CFAR),A*STAR,新加坡)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07461 2025-05-19 cs.SD cs.AI 57%

JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata

Abhinaba Roy, Renhang Liu, Tongyu Lu, Dorien Herremans

机构 * Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.13933 2025-05-16 cs.CV 57%

Unsupervised Video Highlight Detection by Learning from Audio and Visual Recurrence

Zahidul Islam, Sujoy Paul, Mrigank Rochan

机构 * University of Saskatchewan(萨斯喀彻温大学) Google DeepMind(谷歌DeepMind)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Accepted to the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03244 2025-05-14 cs.SD eess.AS 57%

SonicRAG : High Fidelity Sound Effects Synthesis Based on Retrival Augmented Generation

Yu-Ren Guo, Wen-Kai Tai

机构 * Dept. of Computer Science and Information Engineering(计算机科学与信息工程系) National Taiwan University of Science and Technology(台湾科技大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16267 2025-05-14 cs.SD cs.LG eess.AS q-bio.QM 57%

A Classification Benchmark for Artificial Intelligence Detection of Laryngeal Cancer from Patient Voice

Mary Paterson, James Moor, Luisa Cutillo

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 16 pages, 6 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15118 2025-05-13 cs.CV cs.SD 57%

Improving Sound Source Localization with Joint Slot Attention on Image and Audio

Inho Kim, Youngkil Song, Jicheol Park, Won Hwa Kim, Suha Kwak

机构 * Dept. of CSE, POSTECH(计算机科学与工程系,POSTECH)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09089 2025-05-13 cs.SD cs.HC eess.AS 57%

Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter

Yi Xiao, Harshit Sharma, Victoria Tumanova, Asif Salekin

机构 * Arizona State University(亚利桑那州立大学) Syracuse University(Syracuse大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments 13 pages, 5 figures, ICCPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05755 2025-05-12 cs.SD cs.LG eess.AS 57%

CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech

Madhurananda Pahar, Fuxiang Tao, Bahman Mirheidari, Nathan Pevy, Rebecca Bright, Swapnil Gadgil, Lise Sproson, Dorota Braun, Caitlin Illingworth, Daniel Blackburn, Heidi Christensen

机构 * School of Computer Science, University of Sheffield, Sheffield, S1 4DP, UK(1 计算机科学学院,谢菲尔德大学,谢菲尔德,S1 4DP,英国) Therapy Box, London, UK(2 Therapy Box,伦敦,英国) NIHR Devices for Dignity HTC, Sheffield Teaching Hospitals NHS Foundation Trust, Sheffield, S10 2JF, UK(3 NIHR尊严设备HTC,谢菲尔德教学医院国家健康服务基金会信托,谢菲尔德,S10 2JF,英国) Sheffield Institute for Translational Neuroscience (SITraN), University of Sheffield, Sheffield, S10 2HQ, UK(4 谢菲尔德转化神经科学研究所(SITraN),谢菲尔德大学,谢菲尔德,S10 2HQ,英国)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments This paper has been accepted for publication in IEEE SSCI 2025. Copyright belongs to IEEE

Journal ref IEEE SSCI, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03033 2025-05-07 cs.AI cs.HC 57%

Evaluating the Impact of AI-Powered Audiovisual Personalization on Learner Emotion, Focus, and Learning Outcomes

George Xi Wang, Jingying Deng, Safinah Ali

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00525 2025-05-02 eess.IV cs.CV cs.LG 57%

A Methodological and Structural Review of Parkinsons Disease Detection Across Diverse Data Modalities

Abu Saleh Musa Miah, taro Suzuki, Jungpil Shin

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04644 2025-05-01 eess.AS cs.SD 57%

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

Hanzhao Li, Yuke Li, Xinsheng Wang, Jingbin Hu, Qicong Xie, Shan Yang, Lei Xie

机构 * Hong Kong University of Science and Technology(香港科技大学) Tencent AI Lab(腾讯AI实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 14 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20016 2025-04-29 cs.HC cs.CY cs.MM 57%

Applying LLM-Powered Virtual Humans to Child Interviews in Child-Centered Design

Linshi Li, Hanlin Cai

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM

Comments This paper has been accepted as a Work-in-Progress (WiP) paper in the 24th annual ACM Interaction Design and Children (IDC) Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01089 2025-04-29 cs.RO cs.AI 57%

HomeEmergency -- Using Audio to Find and Respond to Emergencies in the Home

James F. Mullen, Dhruva Kumar, Xuewei Qi, Rajasimman Madhivanan, Arnie Sen, Dinesh Manocha, Richard Kim

机构 * Amazon Lab126(亚马逊实验室126) University of Maryland(马里兰大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Journal ref IEEE Robotics and Automation Letters (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16573 2025-04-24 cs.HC cs.AI 57%

PsyCounAssist: A Full-Cycle AI-Powered Psychological Counseling Assistant System

Xianghe Liu, Jiaqi Xu, Tao Sun

机构 * Beijing PsychTech Technology Co., Ltd.(北京心理科技技术有限公司) University of Hechi(贺州学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04067 2025-04-24 cs.CV 57%

FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait Synthesis

Ziqi Ni, Ao Fu, Yi Zhou

机构 * School of Computer Science and Engineering, Southeast University(计算机科学与工程学院,东南大学) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Ministry of Education, China(新一代人工智能技术及其交叉应用重点实验室,教育部,中国)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Accepted by ICMR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09269 2025-04-22 cs.SD cs.LG eess.AS 57%

Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs

Anshuman Sinha, Camille Migozzi, Aubin Rey, Chao Zhang

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments 29 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13205 2025-04-21 cs.CR cs.AI 57%

On-Device Watermarking: A Socio-Technical Imperative For Authenticity In The Age of Generative AI

Houssam Kherraz

机构 * Kensho Technologies

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

Comments 10 pages, 3 figures, ICLR 2025, https://openreview.net/forum?id=ygE0U21vxM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06778 2025-04-18 cs.SD eess.AS 57%

CAFA: a Controllable Automatic Foley Artist

Roi Benita, Michael Finkelson, Tavi Halperin, Gleb Sterkin, Yossi Adi

机构 * Technion(以色列理工学院) Hebrew University of Jerusalem(耶路撒冷希伯来大学) Lightricks(丽塔克斯科技公司)

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Renamed paper to "CAFA: a Controllable Automatic Foley Artist" from "Controllable Automatic Foley Artist". Updated link to demo page

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19898 2025-04-16 cs.LG cs.AI 57%

A Review of Deep Learning Approaches for Non-Invasive Cognitive Impairment Detection

Muath Alsuhaibani, Ali Pourramezan Fard, Jian Sun, Farida Far Poor, Peter S. Pressman, Mohammad H. Mahoor

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08687 2025-04-14 cs.HC cs.AI cs.CY 57%

Voice Interaction With Conversational AI Could Facilitate Thoughtful Reflection and Substantive Revision in Writing

Jiho Kim, Philippe Laban, Xiang 'Anthony' Chen, Kenneth C. Arnold

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Comments 5 pages; Accepted to Fourth Workshop on Intelligent and Interactive Writing Assistants (In2Writing 2025) at NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17911 2025-04-08 cs.CV cs.CR 57%

Passive Deepfake Detection Across Multi-modalities: A Comprehensive Survey

Hong-Hanh Nguyen-Le, Van-Tuan Tran, Dinh-Thuc Nguyen, Nhien-An Le-Khac

机构 * University College Dublin(都柏林大学学院) Trinity College Dublin(都柏林圣三一学院) University of Science(科学大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments 35 pages

详情

展开后加载摘要…

URL PDF HTML 收藏