arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4567 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4567 篇

2601.13143 2026-01-21 cs.LG 82%

FastAV: Efficient Token Pruning for Audio-Visual Large Language Model Inference

FastAV: 面向音频视觉大语言模型推理的高效令牌修剪

Chaeyoung Jung, Youngjoon Jang, Seungwoo Lee, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract)

AI总结 FastAV通过两阶段令牌修剪策略,针对音频视觉大语言模型实现高效推理,减少FLOPs超过40%并保持性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13098 2026-01-21 cs.HC 82%

Exploring the Impacts of Background Noise on Auditory Stimuli of Audio-Visual eHMIs for Hearing, Deaf, and Hard-of-Hearing People

探索背景噪声对听觉刺激在音频视觉eHMIs中的影响:面向听障、聋人及听力障碍人群

Wenge Xu, Foroogh Hajiseyedjavadi, Debargha Dey, Tram Thi Minh Tran, Mark Colley

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract)

AI总结 研究探讨了背景噪声对聋人及听力障碍者在音频视觉eHMI中听觉刺激的影响,发现嘈杂环境会损害行人穿越体验,而增加铃声或语音提示可改善体验。

Comments This is the author's version of the paper accepted at CHI Conference on Human Factors in Computing Systems (CHI '26), April 13-17, 2026, Barcelona, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05879 2026-01-19 cs.HC 82%

Human-AI Alignment of Multimodal Large Language Models with Speech-Language Pathologists in Parent-Child Interactions

与语言病理学家在亲子互动中的人机对齐多模态大语言模型

Weiyan Shi, Kenny Tsu Wei Choo

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract)

AI总结 本研究通过多模态大语言模型与语言病理学家的对齐,开发了支持亲子互动分析的系统,实现了85%的感知线索提取准确率和75%的判断精确率,并提出了行为观察-判断系统的构建指南。

Comments This is an earlier version of the work released in May 2025. The version accepted at CHI 2026 is available as a separate preprint at arXiv:2511.04366

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04781 2026-01-09 cs.HC 82%

Dynamic Thermal Feedback in Highly Immersive VR Scenarios: a Multimodal Analysis of User Experience

高度沉浸式VR场景中的动态热反馈:用户体验的多模态分析

Sophie Villenave, Pierre Raimbaud, Guillaume Lavoué

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract)

AI总结 本研究通过多模态分析探讨了高度沉浸式VR场景中动态热反馈对用户体验的影响,发现热反馈可提升沉浸感,但不同质量水平对用户体验的影响因人而异。

Comments 15 pages, 9 figures. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13673 2025-12-30 cs.MM cs.CV cs.IR cs.LG eess.AS 82%

Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

深度音频视觉相关性学习的最新进展与挑战

Luís Vilaça, Yi Yu, Paula Viana

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

AI总结 本文综述了深度音频-视觉相关性学习的最新进展,分析了现有模型、优化方法及未来研究方向。

Comments 8 pages, 1 figure

Journal ref ACM Computing Surveys, 57(12), 1-46, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14115 2025-12-17 cs.SD cs.LG 82%

Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting

多模态对比学习联合框架用于鲁棒的语音术语检测和关键词 spotting

Ramesh Gundluru, Shubham Gupta, Sri Rama Murty K

机构 * Electrical Engineering(电子工程) Artificial Intelligence(人工智能)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

AI总结 本文提出一种联合多模态对比学习框架,通过统一音频和跨模态监督提升语音术语检测和关键词 spotting 的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04720 2025-12-05 cs.SD 82%

M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis

M3-TTS: 多模态DiT对齐与梅尔-潜在空间用于零样本高质量语音合成

Xiaopeng Wang, Chunyu Qiang, Ruibo Fu, Zhengqi Wen, Xuefei Liu, Yukun Liu, Yuzhe Liang, Kang Yin, Yuankun Xie, Heng Xie, Chenxing Li, Chen Zhang, Changsheng Li

机构 * Beijing Institute of Technology(北京理工大学) Kuaishou Technology(快手科技) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 音频语音多模态 :multi-modal(title,abstract);cross-modal(abstract)

AI总结 M3-TTS通过多模态扩散变换器实现零样本高质量语音合成,采用联合扩散层和梅尔-变体编码器,实现高效且自然的语音生成。

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17323 2025-11-24 cs.SD cs.AI cs.CL cs.MM 82%

MusicAIR: A Multimodal AI Music Generation Framework Powered by an Algorithm-Driven Core

MusicAIR: 一种由算法驱动核心的多模态AI音乐生成框架

Callie C. Liao, Duoduo Liao, Ellie L. Zhang

机构 * Stanford University(斯坦福大学) George Mason University(乔治·马歇尔大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

AI总结 MusicAIR通过算法驱动的核心生成音乐,能从歌词、文本和图像生成符合音乐理论的旋律谱,提升音乐创作效率并降低入门门槛。

Comments Accepted by IEEE Big Data 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01094 2025-11-21 cs.SD cs.AI cs.MM eess.AS 82%

MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions

MMVA: 基于估值和唤醒度的多模态匹配

Suhwan Choi, Kyu Won Kim, Myungjoo Kang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM、eess.AS

AI总结 MMVA通过多模态匹配基于估值和唤醒度,实现跨图像、音乐和音乐字幕的情感内容捕捉,并在估值-唤醒度预测任务中取得最佳性能。

Comments Paper accepted in Artificial Intelligence for Music workshop at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10069 2025-11-12 cs.DC cs.LG 82%

ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism

Zedong Liu, Shenggan Cheng, Guangming Tan, Yang You, Dingwen Tao

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) University of Electronic Science and Technology of China(电子科技大学) National University of Singapore(新加坡国立大学)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract)

Comments Accepted at NeurIPS 2025 Oral (Thirty-Ninth Conference on Neural Information Processing Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19281 2025-11-12 cs.RO 82%

Audio-Visual Traffic Light State Detection for Urban Robots

Sagar Gupta, Akansel Cosgun

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);multi-modal(abstract)

Comments Submitted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2024

Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06988 2025-11-11 cs.LG cs.HC 82%

HCFSLN: Adaptive Hyperbolic Few-Shot Learning for Multimodal Anxiety Detection

Aditya Sneh, Nilesh Kumar Sahu, Anushka Sanjay Shelke, Arya Adyasha, Haroon R. Lone

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06288 2025-11-11 cs.SD cs.CL cs.MM eess.AS 82%

ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction

Wenxuan Wu, Shuai Wang, Xixin Wu, Helen Meng, Haizhou Li

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CL、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13145 2025-11-03 cs.SD 82%

UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model

Yudong Yang, Xiaokang Liu, Shaofeng zhao, Rongfeng Su, Nan Yan, Lan Wang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, China(深圳先进技术研究院,中国科学院,中国) Key Laboratory of Biomedical Imaging Science and System, Chinese Academy of Sciences, China(生物医学成像科学与系统重点实验室,中国科学院,中国)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17113 2025-10-28 cs.CV cs.AI cs.CL 82%

MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation

Shoubin Yu, Yue Zhang, Ziyang Wang, Jaehong Yoon, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Nanyang Technological University(南洋理工大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments EMNLP 2025 Findings; The first two authors contributed equally; Github link: https://github.com/Yui010206/MEXA

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02236 2025-10-22 cs.CV cs.MM cs.SD eess.AS 82%

3D Audio-Visual Segmentation

Artem Sokolov, Swapnil Bhosale, Xiatian Zhu

机构 * University of Surrey, UK(Surrey大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at the NeurIPS 2024 Workshop on Audio Imagination; this version updates the project page link

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13182 2025-10-16 cs.LG 82%

Information-Theoretic Criteria for Knowledge Distillation in Multimodal Learning

Rongrong Xie, Yizhou Xu, Guido Sanguinetti

机构 * Scuola Internazionale Superiore di Studi Avanzati (SISSA)(国际先进研究高等学院) École Polytechnique Fédérale de Lausanne (EPFL)(日内瓦联邦理工学院)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06060 2025-10-08 cs.MM cs.AI cs.CV 82%

Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information

Christian Marinoni, Riccardo Fosco Gramaccioni, Eleonora Grassucci, Danilo Comminiello

机构 * Sapienza University of Rome, Italy(罗马大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02165 2025-10-03 cs.SE 82%

Towards fairer public transit: Real-time tensor-based multimodal fare evasion and fraud detection

Peter Wauyo, Dalia Bwiza, Alain Murara, Edwin Mugume, Eric Umuhoza

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01690 2025-10-03 cs.GR cs.HC 82%

Multimodal Feedback for Task Guidance in Augmented Reality

Hu Guo, Lily Patel, Rohan Gupt

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15017 2025-09-30 cs.CL cs.AI cs.SD eess.AS 82%

DM-Codec: Distilling Multimodal Representations for Speech Tokenization

Md Mubtasim Ahasan, Md Fahim, Tasnim Mohiuddin, A K M Mahbubur Rahman, Aman Chadha, Tariq Iqbal, M Ashraful Amin, Md Mofijul Islam, Amin Ahsan Ali

机构 * Center for Computational & Data Sciences, Independent University, Bangladesh(计算与数据科学中心,独立大学,孟加拉国) Amazon GenAI(亚马逊生成人工智能) Qatar Computing Research Institute(卡塔尔计算研究所) University of Virginia(弗吉尼亚大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23833 2025-09-30 eess.AS cs.CV cs.MM cs.SD 82%

AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines

Cancan Li, Fei Su, Juan Liu, Hui Bu, Yulong Wan, Hongbin Suo, Ming Li

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Beijing AISHELL Technology Co., Ltd.(北京AISHELL科技有限公司) AI Center, OPPO(OPPO人工智能中心)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20724 2025-09-26 cs.SI cs.CL cs.CV cs.MM 82%

Visual Authority and the Rhetoric of Health Misinformation: A Multimodal Analysis of Social Media Videos

Mohammad Reza Zarei, Barbara Stead-Coyle, Michael Christensen, Sarah Everts, Majid Komeili

机构 * School of Computer Science(计算机科学学院) Carleton University(卡尔顿大学) Department of Law and Legal Studies(法律与法律研究系) School of Journalism and Communication(新闻与传播学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19631 2025-09-25 eess.AS cs.AI cs.CL 82%

Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning

Shaoshi Ling, Gang Liu, Guoli Ye, Jinyu Li

机构 * Microsoft CoreAI(微软核心人工智能)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18816 2025-09-24 cs.SD cs.CL cs.MM eess.AS 82%

Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models

Junyu Wang, Ziyang Ma, Zhengding Luo, Tianrui Wang, Meng Ge, Xiaobao Wang, Longbiao Wang

机构 * 1Laboratory of Cognitive Computing Application, College of Intelligence Computing, Tianjin University, Tianjin, China 2School of Computer Science, Shanghai Jiao Tong University, Shanghai, China 3School of Electrical \& Electronic Engineering, Nanyang Technological University, Singapore 4Huiyan Technology (Tianjin) Co., Ltd, Tianjin, China

专题命中 音频语音多模态 :cross-modal(title);multi-modal(abstract);分类 cs.CL、cs.MM、eess.AS

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02823 2025-09-01 cs.SD cs.AI cs.MM eess.AS 82%

A Multimodal Symphony: Integrating Taste and Sound through Generative AI

Matteo Spanio, Massimiliano Zampini, Antonio Rodà, Franco Pierucci

机构 * Centro di Sonologia Computazionale (CSC)(计算声学中心) Department of Information Engineering University of Padova(信息工程系帕多瓦大学) Center for Mind/Brain Sciences (CIMeC)(心智/大脑科学中心) University of Trento(特伦托大学) SoundFood s.r.l.(SoundFood公司)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM、eess.AS

Comments 17 pages, 6 figures (2 + 2 figures with 2 subfigures each)

Journal ref Front. Comput. Sci. 7:1575741 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20359 2025-08-29 cs.IR 82%

Progressive Semantic Residual Quantization for Multimodal-Joint Interest Modeling in Music Recommendation

Shijia Wang, Tianpei Ouyang, Qiang Xiao, Dongjing Wang, Yintao Ren, Songpei Xu, Da Guo, Chuanjiang Luo

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05087 2025-08-08 cs.MM cs.AI cs.CL cs.CR 82%

JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering

Renmiao Chen, Shiyao Cui, Xuancheng Huang, Chengwei Pan, Victor Shea-Jay Huang, QingLin Zhang, Xuan Ouyang, Zhexin Zhang, Hongning Wang, Minlie Huang

机构 * CoAI group, DCST, Tsinghua University(清华大学DCST学院) Beihang University(北航大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments 10 pages, 3 tables, 2 figures, to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23010 2025-08-01 cs.LG cs.AI cs.CV cs.SD eess.AS 82%

Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods

Siwoo Park

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06273 2025-07-22 cs.CV cs.MM cs.SD eess.AS 82%

Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

Jeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis, Yong Man Ro

机构 * KAIST(韩国科学技术院) Imperial College London(伦敦帝国理工学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted at ICCV 2025. Code available at: https://github.com/JeongHun0716/zero-avsr

详情

展开后加载摘要…

URL PDF HTML 收藏