arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46237 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4578 篇

2511.12404 2025-11-18 cs.MM cs.AI cs.SD 81%

SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs

Shail Desai, Aditya Pawar, Li Lin, Xin Wang, Shu Hu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19999 2025-11-05 cs.MM cs.CV cs.SD 81%

MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization

Jianxuan Yang, Xiaoran Yang, Lipan Zhang, Xinyue Guo, Zhao Wang, Gongping Huang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24024 2025-10-29 eess.AS cs.CV eess.IV 81%

Listening without Looking: Modality Bias in Audio-Visual Captioning

Yuchi Ishikawa, Toranosuke Manabe, Tatsuya Komatsu, Yoshimitsu Aoki

机构 * LY Corporation(LY公司) Keio University(庆应大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22455 2025-10-28 cs.SD cs.AI eess.AS 81%

Evaluating Multimodal Large Language Models on Core Music Perception Tasks

Brandon James Carone, Iran R. Roman, Pablo Ripollés

机构 * Department of Psychology, Music and Audio Research Laboratory(心理学系、音乐与音频研究实验室) Department of Electronic Engineering and Computer Science(电子工程与计算机科学系)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、eess.AS

Comments Accepted to the NeurIPS 2025 Workshop on AI for Music (AI4Music), 16 pages, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13979 2025-10-17 cs.AI cs.CL 81%

Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talks

Supriti Sinhamahapatra, Jan Niehues

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13281 2025-10-16 eess.AS cs.CL cs.LG 81%

Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses

Sungnyun Kim, Kangwook Jang, Sungwoo Cho, Joon Son Chung, Hoirin Kim, Se-Young Yun

机构 * Kim Jaechul Graduate School of AI, KAIST(金 Jaechul人工智能研究生院,韩国科学技术院) School of Electrical Engineering, KAIST(电气工程学院,韩国科学技术院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CL、eess.AS

Comments Preprint work

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12133 2025-10-15 cs.CL cs.AI 81%

SafeMT: Multi-turn Safety for Multimodal Language Models

Han Zhu, Juntao Dai, Jiaming Ji, Haoran Li, Chengkun Cai, Pengcheng Wen, Chi-Min Chan, Boyuan Chen, Yaodong Yang, Sirui Han, Yike Guo

机构 * Hong Kong University of Science and Technology(香港理工大学) Peking University(北京大学) University of Edinburgh(爱丁堡大学)

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08004 2025-10-10 cs.SD cs.MM eess.AS 81%

Personality-Enhanced Multimodal Depression Detection in the Elderly

Honghong Wang, Jing Deng, Rong Zheng

机构 * Beijing Fosafer Information Technology Co., Ltd.(北京福萨弗信息科技有限公司)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM、eess.AS

Comments 6 pages,2 figures,accepted by ACM Multimedia Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04136 2025-10-07 eess.AS cs.CV cs.SD 81%

MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition

Umberto Cappellazzo, Minsu Kim, Pingchuan Ma, Honglie Chen, Xubo Liu, Stavros Petridis, Maja Pantic

机构 * Imperial College London(伦敦帝国学院) Meta AI NatWest AI Research(NatWest人工智能研究)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02322 2025-10-06 eess.AS cs.CL 81%

SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis

Lukas Buess, Jan Geier, David Bani-Harouni, Chantal Pellegrini, Matthias Keicher, Paula Andrea Perez-Toro, Nassir Navab, Andreas Maier, Tomas Arias-Vergara

机构 * Computer Aided Medical Procedures, Technical University of Munich(慕尼黑技术大学计算机辅助医学程序)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、eess.AS

Comments Submitted to ICASSP 2026; under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00485 2025-10-02 cs.SD cs.AI eess.AS 81%

PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation

Yujia Xiao, Liumeng Xue, Lei He, Xinyi Chen, Aemon Yat Fei Chiu, Wenjie Tian, Shaofei Zhang, Qiuqiang Kong, Xinfa Zhu, Wei Xue, Tan Lee

机构 * The Chinese University of Hong Kong, Hong Kong, China(香港中文大学) The Hong Kong University of Science and Technology, Hong Kong, China(香港科学与技术大学) Microsoft, China(微软公司) South China University of Technology, China(华南理工大学) Northwestern Polytechnical University, China(西北工业大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22718 2025-09-30 eess.AS cs.MM cs.SD 81%

PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos

Ke Gu, Zhicong Wu, Peng Bai, Sitong Qiao, Zhiqi Jiang, Junchen Lu, Xiaodong Shi, Xinyuan Qian

机构 * Xiamen University(厦门大学) University of Science and Technology Beijing(北京科技大学) National University of Singapore(新加坡国立大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10859 2025-09-29 cs.MM cs.CL cs.HC 81%

MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions

Ramaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha

机构 * University of Maryland, College Park(马里兰大学 College Park 分校)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15220 2025-09-29 cs.CV cs.CL cs.SD 81%

video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models

Changli Tang, Yixuan Li, Yudong Yang, Jimin Zhuang, Guangzhi Sun, Wei Li, Zejun Ma, Chao Zhang

机构 * Tsinghua University(清华大学) University of Cambridge(剑桥大学) ByteDance(字节跳动)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18200 2025-09-24 cs.LG cs.AI cs.CL cs.RO 81%

Conversational Orientation Reasoning: Egocentric-to-Allocentric Navigation with Multimodal Chain-of-Thought

Yu Ti Huang

机构 * Trans-disciplinary Bachelor Degree Program National Taiwan University(台湾国立大学跨学科学士学位计划)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14930 2025-09-19 cs.CL cs.AI 81%

Cross-Modal Knowledge Distillation for Speech Large Language Models

Enzhi Wang, Qicheng Li, Zhiyuan Tang, Yuhang Jia

机构 * TMCC, College of Computer Science, Nankai University, Tianjin, China(TMCC,计算机科学学院,南开大学,天津,中国) Tencent Ethereal Audio Lab, Tencent Corporation, Shenzhen, China(腾讯虚实音频实验室,腾讯公司,深圳,中国)

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14627 2025-09-19 cs.HC cs.AI cs.CL 81%

Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech

Taesoo Kim, Yongsik Jo, Hyunmin Song, Taehwan Kim

机构 * Artificial Intelligence Graduate School(人工智能研究生院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Published in Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20805 2025-08-29 cs.CL cs.AI cs.SD 81%

Exploring Machine Learning and Language Models for Multimodal Depression Detection

Javier Si Zhao Hong, Timothy Zoe Delaya, Sherwyn Chan Yin Kit, Pai Chet Ng, Xiaoxiao Miao

机构 * Singapore Institute of Technology(新加坡理工学院) Duke Kunshan University(杜克-昆山大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted by APCIPA ASC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18653 2025-08-27 cs.LG cs.AI cs.SD eess.AS 81%

The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability

Xiaoliang Chen, Xin Yu, Le Chang, Teng Jing, Jiashuai He, Ze Wang, Yangjun Luo, Xingyu Chen, Jiayue Liang, Yuchen Wang, Jiaying Xie

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、eess.AS

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07337 2025-08-12 eess.AS cs.CV 81%

KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features

Ivan Kukanov, Jun Wah Ng

机构 * KLASS Engineering and Solutions Singapore(KLASS工程与解决方案新加坡)

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 cs.CV、eess.AS

Comments 7 pages, accepted to the 33rd ACM International Conference on Multimedia (MM'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04353 2025-08-07 cs.MM cs.AI 81%

LUST: A Multi-Modal Framework with Hierarchical LLM-based Scoring for Learned Thematic Significance Tracking in Multimedia Content

Anderson de Lima Luiz

机构 * AImotion Bavaria Technische Hochschule Ingolstadt(巴伐利亚AImotion技术大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI、cs.MM

Comments 5 pages and 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15641 2025-08-07 cs.CL cs.AI 81%

Leveraging Context for Multimodal Fallacy Classification in Political Debates

Alessio Pittiglio

机构 * DISI, University of Bologna(DISI,博洛尼亚大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments 12th Workshop on Argument Mining (ArgMining 2025) @ ACL 2025

Journal ref In Proceedings of the 12th Argument mining Workshop (ArgMining 2025), pages 388-397, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02786 2025-08-07 cs.SD cs.CV eess.AS 81%

CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation

Yuanhong Chen, Kazuki Shimada, Christian Simon, Yukara Ikemiya, Takashi Shibuya, Yuki Mitsufuji

机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) Sony Group Corporation(索尼集团) Sony AI, Sony Group Corporation(索尼人工智能,索尼集团)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02905 2025-08-06 cs.CV cs.SD eess.AS 81%

How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes

Mahnoor Fatima Saad, Ziad Al-Halah

机构 * University of Utah(犹他大学)

专题命中 音频语音多模态 :multimodal(title);audio-visual(abstract);分类 cs.CV、eess.AS

Comments Accepted to ICCV 2025. Project Page: https://mahnoor-fatima-saad.github.io/m-capa.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02000 2025-08-05 cs.SD cs.CV eess.AS eess.IV 81%

Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling

Xuanjun Chen, Shih-Peng Cheng, Jiawei Du, Lin Zhang, Xiaoxiao Miao, Chung-Che Wang, Haibin Wu, Hung-yi Lee, Jyh-Shing Roger Jang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00760 2025-08-04 cs.CL cs.AI 81%

MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations

Qiyao Xue, Yuchen Dou, Ryan Shi, Xiang Lorraine Li, Wei Gao

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22934 2025-08-01 cs.CL cs.AI 81%

Deep Learning Approaches for Multimodal Intent Recognition: A Survey

Jingwei Zhao, Yuhua Wen, Qifei Li, Minchi Hu, Yingying Zhou, Jingyao Xue, Junyang Wu, Yingming Gao, Zhengqi Wen, Jianhua Tao, Ya Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Tsinghua University(清华大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

Comments Submitted to ACM Computing Surveys

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19956 2025-07-29 cs.CV cs.AI q-bio.NC 81%

Predicting Brain Responses To Natural Movies With Multimodal LLMs

Cesar Kadir Torrico Villanueva, Jiaxin Cindy Tu, Mihir Tripathy, Connor Lane, Rishab Iyer, Paul S. Scotti

机构 * Medical AI Research Center (MedARC)(医学人工智能研究中心(MedARC)) Psychological and Brain Sciences, Dartmouth College(心理学与脑科学系,达特茅斯学院) Core for Advanced Magnetic Resonance Imaging (CAMRI), Baylor College of Medicine(先进磁共振成像核心(CAMRI),贝勒医学院) Sophont Princeton Neuroscience Institute(普林斯顿神经科学研究所)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Code available at https://github.com/MedARC-AI/algonauts2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08714 2025-07-29 cs.CV cs.AI 81%

Versatile Multimodal Controls for Expressive Talking Human Animation

Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li, Zixin Zhu, Sanping Zhou, Ming Yang, Le Wang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi'an Jiaotong University, Ant Group(人机混合增强智能国家级实验室,西安交通大学,蚂蚁集团) National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi'an Jiaotong University(人机混合增强智能国家级实验室,西安交通大学) University at Buffalo(布法罗大学)

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM MM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23822 2025-07-24 cs.CL cs.MM 81%

Speech as a Multimodal Digital Phenotype for Multi-Task LLM-based Mental Health Prediction

Mai Ali, Christopher Lucasius, Tanmay P. Patel, Madison Aitken, Jacob Vorstman, Peter Szatmari, Marco Battaglia, Deepa Kundur

机构 * The Edward S. Rogers Sr. Department of Electrical and Computer Engineering, University of Toronto, Toronto, Canada(电气与计算机工程系,多伦多大学) Division of Engineering Science, University of Toronto, Toronto, Canada(工程科学系,多伦多大学) Cundill Centre for Child and Youth Depression, Centre for Addiction and Mental Health, Toronto, Canada(儿童与青少年抑郁研究中心,成瘾与心理健康中心) Department of Psychology, York University, Toronto, Canada(心理学系,约克大学) The Hospital for Sick Children, Toronto, ON, Canada(多伦多儿童医院) Department of Psychiatry, University of Toronto, Toronto, Canada(精神病学系,多伦多大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.MM

Comments 6 pages, 1 figure, 3 tables. The corresponding author is Mai Ali (maia dot ali at mail dot utoronto dot ca). Christopher Lucasius and Tanmay P. Patel contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏