arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2108.11996 2021-08-30 cs.CV 70%

Drop-DTW: Aligning Common Signal Between Sequences While Dropping Outliers

Nikita Dvornik, Isma Hadji, Konstantinos G. Derpanis, Animesh Garg, Allan D. Jepson

专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.01149 2021-06-03 cs.SD cs.IR eess.AS 70%

Exploring modality-agnostic representations for music classification

Ho-Hsiang Wu, Magdalena Fuentes, Juan P. Bello

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.02656 2021-04-07 cs.CV cs.AI cs.GR cs.MM cs.SD eess.AS eess.IV 70%

Collaborative Learning to Generate Audio-Video Jointly

Vinod K Kurmi, Vipul Bajaj, Badri N Patro, K S Venkatesh, Vinay P Namboodiri, Preethi Jyothi

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments ICASSP 2021 (Accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.08468 2021-04-06 cs.CV cs.SD 70%

Beyond Image to Depth: Improving Depth Prediction using Echoes

Kranti Kumar Parida, Siddharth Srivastava, Gaurav Sharma

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CV

Comments To appear in CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.13716 2021-03-26 cs.CV 70%

Vectorization and Rasterization: Self-Supervised Learning for Sketch and Handwriting

Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Yongxin Yang, Timothy M. Hospedales, Tao Xiang, Yi-Zhe Song

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2021 Code : https://github.com/AyanKumarBhunia/Self-Supervised-Learning-for-Sketch

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.06607 2020-08-18 cs.CV 70%

Self-supervised Contrastive Video-Speech Representation Learning for Ultrasound

Jianbo Jiao, Yifan Cai, Mohammad Alsharid, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble

专题命中 音频语音多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments MICCAI 2020 (early acceptance)

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.15815 2020-08-03 cs.CV cs.HC cs.LG 70%

Looking At The Body: Automatic Analysis of Body Gestures and Self-Adaptors in Psychological Distress

Weizhe Lin, Indigo Orton, Qingbiao Li, Gabriela Pavarini, Marwa Mahmoud

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.08250 2020-04-20 eess.AS cs.LG 70%

How to Teach DNNs to Pay Attention to the Visual Modality in Speech Recognition

George Sterpu, Christian Saam, Naomi Harte

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 eess.AS

Comments in IEEE/ACM Transactions on Audio, Speech, and Language Processing (to appear)

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.10132 2019-12-27 cs.CL 70%

Exploring Context, Attention and Audio Features for Audio Visual Scene-Aware Dialog

Shachi H Kumar, Eda Okur, Saurav Sahay, Jonathan Huang, Lama Nachman

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CL

Comments Presented at the Visual Question Answering and Dialog Workshop, CVPR 2019, Long Beach, USA. arXiv admin note: substantial text overlap with arXiv:1912.10131

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.05894 2019-11-15 cs.SD eess.AS stat.ML 70%

Coincidence, Categorization, and Consolidation: Learning to Recognize Sounds with Minimal Supervision

Aren Jansen, Daniel P. W. Ellis, Shawn Hershey, R. Channing Moore, Manoj Plakal, Ashok C. Popat, Rif A. Saurous

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 eess.AS

Comments This extended version of a ICASSP 2020 submission under same title has an added figure and additional discussion for easier consumption

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.00424 2019-10-02 cs.SD cs.LG eess.AS 70%

AV Speech Enhancement Challenge using a Real Noisy Corpus

Mandar Gogate, Ahsan Adeel, Kia Dashtipour, Peter Derleth, Amir Hussain

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract);分类 eess.AS

Comments arXiv admin note: substantial text overlap with arXiv:1909.10407

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.03760 2019-09-24 eess.AS cs.SD 70%

Time Domain Audio Visual Speech Separation

Jian Wu, Yong Xu, Shi-Xiong Zhang, Lian-Wu Chen, Meng Yu, Lei Xie, Dong Yu

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract);分类 eess.AS

Comments Accepted to ASRU 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.04204 2019-04-10 cs.CV 70%

2.5D Visual Sound

Ruohan Gao, Kristen Grauman

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract);分类 cs.CV

Comments Published in CVPR 2019, project page: http://vision.cs.utexas.edu/projects/2.5D_visual_sound/

详情

展开后加载摘要…

URL PDF HTML 收藏
1812.08407 2018-12-21 cs.CL 70%

Context, Attention and Audio Feature Explorations for Audio Visual Scene-Aware Dialog

Shachi H Kumar, Eda Okur, Saurav Sahay, Juan Jose Alvarado Leanos, Jonathan Huang, Lama Nachman

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract);分类 cs.CL

Comments 7 pages, 2 figures, DSTC7 workshop at AAAI 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.00489 2017-12-08 cs.CL cs.AI cs.CV cs.LG eess.AS 70%

Visual Features for Context-Aware Speech Recognition

Abhinav Gupta, Yajie Miao, Leonardo Neves, Florian Metze

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 5 pages and 3 figures

Journal ref IEEE Xplore (ICASSP) (2017) 5020-5024

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.06986 2016-11-22 cs.CL cs.LG cs.SD 70%

Robust end-to-end deep audiovisual speech recognition

Ramon Sanabria, Florian Metze, Fernando De La Torre

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.01966 2022-11-04 cs.CV cs.MM cs.SD eess.AS eess.IV 69%

MarginNCE: Robust Sound Localization with a Negative Margin

Sooyoung Park, Arda Senocak, Joon Son Chung

专题命中 音频语音多模态 :audio-visual(abstract,comments);分类 cs.CV、cs.MM、eess.AS

Comments Submitted to ICASSP 2023. SOTA performance in Audio-Visual Sound Localization. 5 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.06304 2021-08-23 cs.AI cs.CL cs.CV 69%

What is Multimodality?

Letitia Parcalabescu, Nils Trost, Anette Frank

专题命中 音频语音多模态 :multimodal(abstract,journal_ref);分类 cs.CV、cs.CL、cs.AI

Comments Paper accepted for publication at MMSR 2021; 10 pages, 5 figures

Journal ref Proceedings of the 1st Workshop on Multimodal Semantic Representations (MMSR), 2021, Groningen, Netherlands (Online), Association for Computational Linguistics, p. 1--10

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.27176 2026-08-28 cs.CL cs.AI cs.LG eess.AS 新提交 67%

When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

当文本误导时:面向音频接地对话的不一致感知推理

Yen-Ju Lu, Yuzhe Wang, Yaohan Guan, Xiluo He, Jiarui Hai, Mingrui Liang, Kaavya Chaparala, Thomas Thebaud, Laureano Moro-Velazquez, Najim Dehak, Jesus Villalba

机构 * Center for Language and Speech Processing, Johns Hopkins University(约翰霍普金斯大学语言与语音处理中心)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 本研究针对口语对话理解中基于转录本的捷径问题,构建了含501个问题的受控基准ContraTalk,提出Audio Twin智能体式推理框架,可提升冲突问答案例的准确率并减少文本偏向陷阱选择。

Comments 24 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20719 2026-08-25 cs.SD cs.AI cs.MM eess.AS 版本更新 67%

ONOTE: Hypergraph-Grounded Omnimodal Reasoning for Computational Music Science

ONOTE:面向专家级音乐智能的多模态记谱处理基准测试

Menghe Ma, Siqing Wei, Yuecheng Xing, Ziyue Zhu, Zhenghong Lin, Yaheng Wang, Fanhong Meng, Peijun Han, Luu Anh Tuan, Haoran Luo

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanyang Technological University(南洋理工大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、cs.MM、eess.AS

AI总结 本文提出ONOTE基准测试,通过确定性流程消除记谱系统偏差,揭示多模态模型在感知准确性和音乐理论理解间的根本分歧。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19192 2026-08-05 cs.DC 版本更新 67%

MUSE: A Heterogeneity-Aware Multimedia Search Engine for Mobile SoCs

AME:一种高效的异构代理记忆引擎用于智能手机

Xinkui Zhao, Qingyu Ma, Yifan Zhang, Hengxuan Lou, Sai Liu, Chang Liu, Guanjie Cheng, Naibo Wang, Yueshen Xu

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract)

AI总结 AME是一种为智能手机设计的高效异构代理记忆引擎,通过硬件感知的矩阵流水线和调度方案提升内存处理效率,实现更高吞吐量和更低延迟。

Comments Accepted by the 34th ACM International Conference on Multimedia (MM '26). 9 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20266 2026-08-04 cs.SD 版本更新 67%

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

大型音频语言模型综述:通用性、可信度与展望

Kaiwen Luo, Zhenhong Zhou, Leyan Wang, Liang Lin, Tianyu Shao, Yuanhe Zhang, Yang Xiao, Yuxuan Li, Miao Yu, Kailin Lyu, Jiaming Zhang, Li Sun, Songze Li, Yueming Wu, Ting Dang, Xiaojun Jia, Dongrui Liu, Kai Li, Rohan Kumar Das, Siyuan Liang, Xinfeng Li, Qiankun Li, Jing Chen, Xingjun Ma, Kun Wang, Junhao Dong, Deqing Zou, Yu Cheng, Xia Hu, Zhigang Zeng, Sen Su, Yang Liu, Yu-Gang Jiang, Philip S. Yu, Yew-Soon Ong

机构 * Nanyang Technological University(南洋理工大学) Independent Researcher(独立研究者) The University of Melbourne(墨尔本大学) North China Electric Power University(华北电力大学) Beijing University of Posts and Telecommunications(北京邮电大学) University of Chinese Academy of Sciences(中国科学院大学) University of Science and Technology of China(中国科学技术大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Shanghai AI Laboratory(上海人工智能实验室) Huazhong University of Science and Technology(华中科技大学) Tsinghua University(清华大学) Fortemedia Singapore(富媒体新加坡) Tencent(腾讯) Fudan University(复旦大学) Wuhan University(武汉大学) Chinese University of Hong Kong(香港中文大学) Chongqing University of Posts and Telecommunications(重庆邮电大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract)

AI总结 本文综述了大型音频语言模型的通用性、可信度及未来发展方向,探讨了其架构创新、对齐算法及安全风险,并提出了防御深入、因果音频世界建模等策略以提升音频智能的可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12290 2026-07-15 eess.AS cs.AI cs.CL cs.LG cs.SD 新提交 67%

The Sound of Absence: Audio-Language Embedding Models Struggle with Negation

缺失之音:音频-语言嵌入模型在处理否定时存在困难

Chun-Yi Kuan, Hung-yi Lee

机构 * Graduate Institute of Communication Engineering, National Taiwan University, Taiwan(国家台湾大学通讯工程研究所) Artificial Intelligence Center of Research Excellence (AI-CoRE), National Taiwan University, Taiwan(国家台湾大学人工智能卓越研究中心)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 研究音频-语言嵌入模型在处理否定时的问题,引入NegEval-Audio框架将数据集转换为否定感知任务,发现模型在否定情况下性能大幅下降,肯定偏差是缺陷,需明确否定感知训练目标。

Comments Manuscript in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11120 2026-07-14 cs.CV cs.CL eess.AS 新提交 67%

Simple Features and Honest Calibration for Ambivalence and Hesitancy Recognition in Video

视频中矛盾与犹豫识别的简单特征与诚实校准

Vikas Kumar, Aditya Mishra, Haroon R. Lone

机构 * Indian Institute of Science Education and Research Bhopal(印度科学教育与研究学院博帕尔分院)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.CL、eess.AS

AI总结 针对ABAW 2026 BAH挑战赛中的矛盾与犹豫识别问题,系统结合情感专用多模态表示与语言犹豫线索,经AMF融合及AP加权集成。引入‘ASR擦除时间’构建特征,实验表明语言是最强通道,校准比架构重要,固定阈值AP加权测试集得分0.731。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22378 2026-07-14 cs.SD cs.AI cs.MM eess.AS 67%

Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach

零努力图像到音乐生成:一种可解释的基于RAG的视觉语言模型方法

Zijian Zhao, Dian Jin, Zijing Zhou

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) The Hong Kong Polytechnic University(香港理工大学) The University of Hong Kong(香港大学) The Hong Kong University of Science(香港科学大学) The Hong Kong Polytechnic University Hong Kong China(香港理工大学香港中国) The University of Hong Kong Hong Kong China(香港大学香港中国)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、cs.MM、eess.AS

AI总结 本文提出一种基于RAG的视觉语言模型方法,实现图像到音乐生成,通过ABC记谱法连接文本与音乐模态,并利用多模态检索增强生成和自反思技术,提供高可解释性且低计算成本的音乐生成方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05196 2026-07-08 cs.CL cs.AI cs.LG cs.SD eess.AS 新提交 67%

Unified Audio Intelligence Without Regressing on Text Intelligence

不退化文本智能的统一音频智能

Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim, Boxin Wang, Zihan Liu, Sungwon Kim, Yang Chen, Arushi Goel, Rajarshi Roy, Wenliang Dai, Zhuolin Yang, Yangyi Chen, Dongfu Jiang, Sreyan Ghosh, Tuomas Rintamaki, Andrew Tao, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping

机构 * Nemotron-Labs(Nemotron实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 本文提出统一音频-文本大语言模型Audex,通过单Transformer解码器架构与多阶段训练,在实现多音频任务SOTA性能的同时,几乎不损失骨干文本LLM的原有文本智能能力。

Comments We release the Audex models at https://huggingface.co/collections/nvidia/nemotron-labs-audex

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02504 2026-07-03 cs.CL cs.AI cs.CV 新提交 67%

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

推理大模型提升长篇电视剧中的说话人识别

Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu, Jiacheng Shao, Pengfei Chen, Jiannan Ge, Kaiwen Duan, Qi Tian

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 提出DramaSR-532K大规模基准和基于推理大模型的DramaSR-LRM方法,通过多模态工具使用聚合上下文证据,显著提升长篇电视剧中说话人识别准确率,尤其对短话语效果显著。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09803 2026-07-01 cs.SD 版本更新 67%

MAGE: Modality-Agnostic Music Generation and Target-Source Extraction

MAGE:模态无关的音乐生成与编辑

Muhammad Usama Saleem, Tejasvi Ravi, Tianyu Xu, Rajeev Nongpiur, Ishan Chatterjee, Mayur Jagdishbhai Patel, Pu Wang

机构 * Google(谷歌) University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)

专题命中 音频语音多模态 :multimodal(abstract);audio-visual(abstract)

AI总结 MAGE提出了一种模态无关框架,统一了多模态音乐生成与混合编辑,通过可控多模态FluxFormer和音频视觉 nexus 对齐机制,实现灵活且高效的音乐生成与编辑。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04840 2026-06-24 eess.AS cs.AI cs.CL 版本更新 67%

An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production

一种用于言语产生过程中发音、大脑和肌肉活动的实时MRI视频、脑电图和表面肌电图同步采集方法

Jihwan Lee, Parsa Razmara, Kevin Huang, Sean Foley, Aditya Kommineni, Haley Hsu, Woojae Jeong, Prakash Kumar, Xuan Shi, Yoonjeong Lee, Tiantian Feng, Takfarinas Medani, Ye Tian, Sudarsana Reddy Kadiri, Krishna S. Nayak, Dani Byrd, Louis Goldstein, Richard M. Leahy, Shrikanth Narayanan

机构 * Signal Analysis and Interpretation Laboratory, University of Southern California(南加州大学信号分析与解释实验室) Ming Hsieh Dept. of Electrical and Computer Engineering, University of Southern California(南加州大学明希斯电气与计算机工程系) Dept. of Linguistics, University of Southern California(南加州大学语言学系)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 本文首次实现了言语产生过程中实时MRI、EEG和表面EMG的同步采集,并提出了针对三模态设置的伪影抑制流程,为言语神经科学和脑机接口提供新视角。

Comments Accepted for Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06191 2026-06-23 eess.AS cs.AI cs.CL cs.SD 版本更新 67%

Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment

Harf-Speech:面向阿拉伯语音素级语音评估的临床对齐框架

Asif Azad, MD Sadik Hossain Shanto, Mohammad Sadat Hossain, Bdour Alwuqaysi, Sabri Boughorbel, Yahya Bokhari, Abdulrhman Aljouie, Ayah Othman Sindi, Ehsan Hoque

机构 * Ministry of Defense, Saudi Arabia(沙特阿拉伯国防部) Ability Center, Saudi Arabia(沙特阿拉伯能力中心) University of Rochester, USA(罗切斯特大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 提出Harf-Speech模块化系统,结合MSA音素化器、微调语音到音素模型、Levenshtein对齐及混合评分器,在阿拉伯语音素级发音评估中达到0.791 Pearson相关和0.659 ICC,优于现有端到端方法。

详情

展开后加载摘要…

URL PDF HTML 收藏