arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2603.23673 2026-03-26 eess.AS cs.SD 70%

Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition

Crab:多层对比监督以提升在表演和自然语音条件下的语音情感识别

Lucas H. Ueda, João G. T. Lima, Paula D. P. Costa

机构 * Dept. of Computer Engineering and Automation (DCA), Faculdade de Engenharia Elétrica e de Computação(计算机工程与自动化系(DCA)、电气与计算机工程学院) AI Lab.(人工智能实验室) Institute of Computing, Universidade Estadual de Campinas, UNICAMP(计算学院,坎皮纳斯州立大学,UNICAMP)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 eess.AS

AI总结 本文提出Crab模型,通过多层对比监督策略提升语音情感识别在表演和自然语音条件下的性能,采用跨模态Transformer架构和多任务对比学习,有效解决类别不平衡问题。

Comments IEEE Transactions on Affective Computing submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20894 2026-03-24 cs.MM 70%

AcoustEmo: Open-Vocabulary Emotion Reasoning via Utterance-Aware Acoustic Q-Former

AcoustEmo:通过语句感知的声学Q-Former实现开放词汇情绪推理

Liyun Zhang, Xuanmeng Sha, Shuqiong Wu, Fengkai Liu

专题命中 音频语音多模态 :multimodal(abstract);MLLM(abstract);分类 cs.MM

AI总结 AcoustEmo通过引入语句感知的声学Q-Former,利用时间同步滑动窗口提取段级音频token,提升复杂情绪推理能力,在EMER任务中优于基线模型。

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09212 2026-03-11 eess.AS 70%

Acoustic and Semantic Modeling of Emotion in Spoken Language

语音中情感的声学与语义建模

Soumya Dutta

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 eess.AS

AI总结 本研究通过联合建模声学与语义信息,提升语音中情感的理解与合成能力,提出情感意识的预训练框架和对话场景下的情感识别方法,以及无文本的语音情感风格转换技术。

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05708 2026-03-09 cs.CV 70%

Interpretable Perception and Reasoning for Audiovisual Geolocation

可解释的视听定位与推理

Yiyang Su, Xiaoming Liu

专题命中 音频语音多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出了一种基于可解释感知和多模态推理的视听定位框架,通过合成声学原子与视觉特征,实现高精度全球定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13127 2026-03-09 cs.CV cs.CR 70%

SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge

SPARK:通过协同提示音频与重新上下文化知识实现T2V模型的劫持

Zonghao Ying, Moyang Chen, Nizhang Li, Zhiqiang Wang, Wenxin Zhang, Quanchen Zou, Zonglei Jing, Aishan Liu, Xianglong Liu

机构 * State Key Laboratory of Complex \& Critical Software Environment, Beihang University College of Science, Mathematics Technology, Wenzhou-Kean University 360 AI Security Lab Faculty of Innovation Engineering, Macau University of Science Hong Kong University of Science University of Chinese Academy of Sciences

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 SPARK通过模块化提示设计,利用音频与视觉关联模式,实现对T2V模型的高效劫持,提升攻击成功率23%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01415 2026-03-03 eess.AS 70%

The USTC-NERCSLIP Systems for the CHiME-9 MCoRec Challenge

USTC-NERCSLIP系统参加CHiME-9 MCoRec挑战

Ya Jiang, Ruoyu Wang, Jingxuan Zhang, Jun Du, Yi Han, Zihao Quan, Hang Chen, Yeran Yang, Kongzhi Zheng, Zhuo Chen, Yanhui Tu, Shutong Niu, Changfeng Xi, Mengzhi Wang, Zhongbin Wu, Jieru Chen, Henghui Zhi, Weiyi Shi, Shuhang Wu, Genshun Wan, Jia Pan, Jianqing Gao

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 eess.AS

AI总结 USTC-NERCSLIP系统通过多模态级联方法和LLM技术,在CHiME-9 MCoRec挑战中实现15.70%的联合ASR-聚类错误率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15340 2026-02-27 cs.CV 70%

Towards Seamless Interaction: Causal Turn-Level Modeling of Interactive 3D Conversational Head Dynamics

迈向无缝交互:交互式3D对话头部动力学的因果回合级建模

Junjie Chen, Fei Wang, Zhihao Huang, Qing Zhou, Kun Li, Dan Guo, Linfeng Zhang, Xun Yang

机构 * HFUT(合肥工业大学) IAI, Hefei Comprehensive National Science Center(人工智能研究院,合肥综合性国家科学中心) USTC(中国科学技术大学) SJTU(上海交通大学) TeleAI, China Telecom(TeleAI,中国电信) NWPU(西北工业大学) UAEU(阿联酋大学) AHPU(安徽大学)

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 TIMAR通过因果回合级建模,实现了交互式3D对话头部动态的高效生成,提升了对话的表达性和时序一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13954 2026-02-17 cs.SD cs.AI 70%

Eureka-Audio: Triggering Audio Intelligence in Compact Language Models

Eureka-Audio: 在紧凑语言模型中触发音频智能

Dan Zhang, Yishu Lei, Jing Hu, Shuwei He, Songhe Deng, Xianlong Luo, Danxiang Zhu, Shikun Feng, Rui Liu, Jingzhou He, Yu Sun, Hua Wu, Haifeng Wang

专题命中 音频语音多模态 :cross-modal(abstract);omni-modal(abstract);分类 cs.AI

AI总结 Eureka-Audio通过紧凑架构和统一端到端设计,在低参数量下实现高性能音频理解,匹配甚至超越大模型表现。

Comments 23 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11596 2026-02-13 cs.AI 70%

MAPLE: Modality-Aware Post-training and Learning Ecosystem

MAPLE: 多模态感知的后训练与学习生态系统

Nikhil Verma, Minjung Kim, JooYoung Yoo, Kyung-Min Jin, Manasa Bharadwaj, Kevin Ferreira, Ko Keun Kim, Youngjoon Kim

机构 * LG Electronics Toronto AI Lab, Toronto, Canada(LG电子多伦多AI实验室) LG Electronics CTO AI Lab, Seoul, Republic of Korea(LG电子CTO AI实验室)

专题命中 音频语音多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.AI

AI总结 MAPLE通过模态感知的后训练与学习生态系统,提升多模态强化学习的准确性、收敛速度和稳定性。

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02591 2026-02-04 cs.SD cs.AI 70%

VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis

VividVoice: 一个统一的框架,用于场景感知的视觉驱动语音合成

Chengyuan Ma, Jiawei Jin, Ruijie Xiong, Chunxiang Jin, Canxiang Yan, Wenming Yang

机构 * Shenzhen International Graduate School, Tsinghua University, China(清华大学深圳国际研究生院) Ant Group, China(蚂蚁集团)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 VividVoice通过构建大规模多模态数据集和设计核心对齐模块,提升了场景感知的视觉驱动语音合成的音频保真度和多模态一致性。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19469 2026-01-26 cs.SD cs.MM 70%

MusiCRS: Benchmarking Audio-Centric Conversational Recommendation

MusiCRS:面向音频导向的对话推荐基准测试

Rohan Surana, Amit Namburi, Gagan Mundada, Abhay Lal, Zachary Novack, Julian McAuley, Junda Wu

机构 * University of California, San Diego, USA(加州大学圣地亚哥分校)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.MM

AI总结 MusiCRS通过链接Reddit用户对话与音乐曲目,为音频导向的对话推荐提供首个基准测试,揭示了跨模态整合的局限性,并释放了相关数据集和代码以促进研究进展。

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02753 2026-01-07 eess.AS 70%

Vclip: Face-based Speaker Generation by Face-voice Association Learning

基于面部的语音生成:通过面部-语音关联学习

Yao Shi, Yunfei Xu, Hongbin Suo, Yulong Wan, Haifeng Liu

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 eess.AS

AI总结 Vclip通过面部-语音关联学习实现基于面部的语音生成,利用CLIP编码器和检索策略提升语音合成质量。

Comments work done in 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02005 2025-12-30 cs.CV 70%

Learning Visual Affordance from Audio

从音频学习视觉可及性

Lidong Lu, Guo Chen, Zhu Wei, Yicheng Liu, Tong Lu

机构 * Nanjing University(南京大学) China Mobile Communications Company Limited Research Institute(中国移动通信有限公司研究院)

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 AVAGFormer通过融合音频和视觉信号,实现了从音频学习视觉可及性的任务,提升了交互区域的识别性能。

Comments 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12424 2025-12-19 cs.LG cs.AI cs.IR 70%

Multi-Modality Collaborative Learning for Sentiment Analysis

多模态协作学习用于情感分析

Shanmin Wang, Chengguang Liu, Qingshan Liu

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出多模态协作学习框架,通过解耦模块和策略模型提升跨模态情感特征学习,实验证明在四个数据库上性能显著提升。

Comments The method has flaws, especially with the decoupling module. During the decoupling process, the heterogeneity of the three modal data and the differences in distribution were not taken into account

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12077 2025-12-02 cs.CV 70%

Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound

通过视觉学习听觉:现在是让视觉语言模型理解艺术情感的时候了

Dengming Zhang, Weitao You, Jingxiong Li, Weishen Lin, Wenda Shi, Xue Zhao, Heda Zuo, Junxian Wu, Lingyun Sun

机构 * Zhejiang University(浙江大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 VAEmotionLLM通过两阶段框架,利用有限音频预训练使视觉语言模型理解艺术中的多模态情感

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22025 2025-12-01 cs.CV 70%

Layover or Direct Flight: Rethinking Audio-Guided Image Segmentation

转折或直接飞行:重新思考音频引导的图像分割

Joel Alberto Santos, Zongwei Wu, Xavier Alameda-Pineda, Radu Timofte

机构 * Computer Vision Lab, CAIDAS & IFI, University of Würzburg, Germany(计算机视觉实验室、CAIDAS与IFI、乌尔姆大学、德国) Inria at Univ. Grenoble Alpes, CNRS, LJK, France(Inria于格勒诺布尔阿尔卑斯大学、CNRS、LJK、法国)

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 本文提出通过直接音频-视觉对齐进行图像分割,无需依赖文本转录,提升了鲁棒性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11000 2025-11-18 cs.SD cs.AI 70%

DialogGraph-LLM: Graph-Informed LLMs for End-to-End Audio Dialogue Intent Recognition

HongYu Liu, Junxin Li, Changxi Guo, Hao Chen, Yaqian Huang, Yifu Guo, Huan Yang, Lihua Cai

机构 * South China Normal University, Guangzhou, China(华南师范大学) Xiamen Rekey Medical Technology Co., LTD, Xiamen, China(厦门瑞康医疗科技有限公司)

专题命中 音频语音多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

Comments 8 pages, 2 figures. To appear in: Proceedings of the 28th European Conference on Artificial Intelligence (ECAI 2025), Frontiers in Artificial Intelligence and Applications, Vol. 413. DOI: 10.3233/FAIA251182

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04592 2025-09-30 cs.CV 70%

Face-voice Association in Multilingual Environments (FAME) 2026 Challenge Evaluation Plan

Marta Moscati, Ahmed Abdullah, Muhammad Saad Saeed, Shah Nawaz, Rohan Kumar Das, Muhammad Zaigham Zaheer, Junaid Mir, Muhammad Haroon Yousaf, Khalid Malik, Markus Schedl

机构 * Johannes Kepler University Linz(约翰·凯撒大学林茨分校) National University of Computer and Emerging Sciences(国家计算机与新兴科学大学) University of Michigan(密歇根大学) Fortemedia Singapore(Fortemedia新加坡分公司) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) University of Engineering and Technology Taxila(塔希尔工程与技术大学) Human-centered AI Group(以人为本的人工智能小组)

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

Comments 4 pages, ICASSP'26, SP Grand Challenge'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16922 2025-09-23 cs.SD cs.AI eess.IV 70%

PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density Control

Tianheng Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun, Wendong Zheng

机构 * Xinjiang Multimodal Intelligent Processing and Information Security Engineering Technology Research Center(新疆多模态智能处理与信息安全工程技术创新研究中心) School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) School of Electrical Engineering and Automation, Tianjin University of Technology(天津理工大学电气工程与自动化学院)

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.AI

Comments Main paper (15 pages). Accepted for publication by ICONIP( International Conference on Neural Information Processing) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02318 2025-09-23 cs.SD cs.AI cs.CL cs.LG cs.MM eess.AS 70%

Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models

Zhifei Xie, Mingbao Lin, Zihang Liu, Pengcheng Wu, Shuicheng Yan, Chunyan Miao

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(国立新加坡大学) Rakuten

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments Technical report, in process

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16375 2025-09-23 cs.CL 70%

Whisper-UT: A Unified Translation Framework for Speech and Text

Cihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz, Keith Cunningham, Kenton Murray, Kevin Duh, Luis Tavarez-Arce, Paul McNamee, Sanjeev Khudanpur

机构 * Center for Language and Speech Processing, Johns Hopkins University(语言与语音处理中心,约翰霍普金斯大学) Human Language Technology Center of Excellence, Johns Hopkins University(人类语言技术卓越中心,约翰霍普金斯大学) Georgetown University(乔治城大学)

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CL

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15775 2025-09-22 cs.SD eess.AS 70%

EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model

Yiqing Yang, Man-Wai Mak

机构 * Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(电子与电气工程系,香港理工大学)

专题命中 音频语音多模态 :multimodal(abstract);MLLM(abstract);分类 eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16724 2025-09-19 cs.SD eess.AS 70%

SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing

Jinbo Hu, Yin Cao, Ming Wu, Zhenbo Luo, Jun Yang

机构 * Institute of Acoustics, Chinese Academy of Sciences, China(中国科学院声学研究所) MiLM Plus, Xiaomi Inc., China(小米公司) Xi’an Jiaotong Liverpool University, China(西安交通大学利物浦大学) University of Chinese Academy of Sciences, China(中国科学院大学)

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12204 2025-09-16 cs.CV 70%

Character-Centric Understanding of Animated Movies

Zhongrui Gui, Junyu Xie, Tengda Han, Weidi Xie, Andrew Zisserman

机构 * Visual Geometry Group Dept.\ of Engineering Science University of Oxford, UK School of Artifitial Intelligence\ Jiao Tong University Shanghai China School of Artifitial Intelligence\ Jiao Tong University

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09747 2025-09-15 cs.LG cs.AI cs.RO 70%

D-CAT: Decoupled Cross-Attention Transfer between Sensor Modalities for Unimodal Inference

Leen Daher, Zhaobo Wang, Malcolm Mielle

机构 * Ecole Polytechnique Federale de Lausanne(瑞士联邦理工学院洛桑分校) Schindler EPFL Lab(Schindler EPFL实验室)

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03212 2025-09-04 cs.CV 70%

AIVA: An AI-based Virtual Companion for Emotion-aware Interaction

Chenxi Li

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15716 2025-08-25 cs.HC cs.AI 70%

Foundation Models for Cross-Domain EEG Analysis Application: A Survey

Hongqi Li, Yitong Chen, Yujuan Wang, Weihang Ni, Haodong Zhang

机构 * School of Software, Northwestern Polytechnical University(软件学院,西北工业大学)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments Submitted to IEEE Journals

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14130 2025-08-21 eess.AS cs.LG 70%

EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition

Hugo Thimonier, Antony Perzo, Renaud Seguier

机构 * Emobot CentraleSupélec IETR (UMR CNRS 6164)(IETR)

专题命中 音频语音多模态 :multimodal(abstract);multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21847 2025-08-19 cs.CV cs.SD 70%

Differentiable Room Acoustic Rendering with Multi-View Vision Priors

Derong Jin, Ruohan Gao

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

Comments ICCV 2025 (Oral); Project Page: https://humathe.github.io/avdar/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02929 2025-08-07 cs.HC cs.SD eess.AS 70%

AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design in Augmented Reality

Brandon Woodard, Margarita Geleta, Joseph J. LaViola, Andrea Fanelli, Rhonda Wilson

机构 * Brown University(布朗大学) University of California at Berkeley(加州大学伯克利分校) University of Central Florida(中央佛罗里达大学) Dolby Laboratories(杜比实验室)

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 eess.AS

Comments Updated abstract

详情

展开后加载摘要…

URL PDF HTML 收藏