arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4585 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4585 篇

2511.18405 2025-11-25 cs.AI cs.HC cs.IR 79%

A Multimodal Conversational Agent for Tabular Data Analysis

面向表格数据分析的多模态对话代理

Mohammad Nour Al Awad, Sergey Ivanov, Olga Tikhonova, Ivan Khodnenko

机构 * ITMO University(ITMO大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 Talk2Data是一种多模态对话代理,通过语音和文本交互实现表格数据分析,结合LLM、ASR、代码生成和TTS技术,提供多轮对话和可验证计算。

Comments \c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17955 2025-11-25 cs.CL 79%

MTikGuard System: A Transformer-Based Multimodal System for Child-Safe Content Moderation on TikTok

MTikGuard系统:一种基于Transformer的多模态系统,用于TikTok上的儿童安全内容审核

Dat Thanh Nguyen, Nguyen Hung Lam, Anh Hoang-Thi Nguyen, Trong-Hop Do

机构 * Faculty of Information Science and Engineering, University of Information Technology(信息科学与工程学院,信息科技大学) Faculty of Software Engineering, University of Information Technology(软件工程学院,信息科技大学) Faculty of Computer Science, University of Information Technology(计算机科学学院,信息科技大学) University of Information Technology, Ho Chi Minh City(信息科技大学,胡志明市) Vietnam National University, Ho Chi Minh City(越南国家大学,胡志明市)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 MTikGuard系统通过多模态分类框架和扩展数据集,实现TikTok儿童安全内容审核的高准确率和实时部署。

Comments Accepted at PACLIC39

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17776 2025-11-25 cs.LG cs.MM 79%

PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning

PrismSSL: 一个接口,多种模态;一个多模态自监督学习的单一接口库

Melika Shirian, Kianoosh Vadaei, Kian Majlessi, Audrina Ebrahimi, Arshia Hemmat, Peyman Adibi, Hossein Karshenas

机构 * dept. Computer Engineering University of Isfahan(计算机工程系 哈尔拉克大学) dept. Computer Engineering University of Texas at Dallas(计算机工程系 德克萨斯大学达拉斯分校) dept. Computer Science University of Oxford(计算机科学系 奥克兰大学) dept. Artificial Intelligence University of Isfahan(人工智能系 哈尔拉克大学)

专题命中 音频语音多模态 :multimodal(title);cross-modal(abstract);分类 cs.MM

AI总结 PrismSSL是一个统一多模态自监督学习的单一接口库,提供模块化代码、分布式训练和图形化仪表盘,支持多种模态和方法的扩展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15312 2025-11-20 cs.CV 79%

A Multimodal Transformer Approach for UAV Detection and Aerial Object Recognition Using Radar, Audio, and Video Data

Mauro Larrat, Claudomiro Sales

机构 * Institute of Exact and Natural Sciences(精确与自然科学研究所) Federal University of Pará(巴西亚马逊联邦大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments 23 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05817 2025-11-19 cs.HC cs.MM cs.SD 79%

TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech

Weiyan Shi, Sunaya Upadhyay, Geraldine Quek, Kenny Tsu Wei Choo

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments Accepted at AAAI 2026 Workshop on Creative AI for Live Interactive Performances (CLIP). To be published in Springer CCIS series

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11930 2025-11-18 cs.HC cs.CV cs.LG cs.SD 79%

Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering

Tianyu Xu, Jihan Li, Penghe Zu, Pranav Sahay, Maruchi Kim, Jack Obeng-Marnu, Farley Miller, Xun Qian, Katrina Passarella, Mahitha Rachumalla, Rajeev Nongpiur, D. Shin

机构 * Google(谷歌)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST '25), Article 17, 1-16, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10325 2025-11-14 cs.MM 79%

TMDC: A Two-Stage Modality Denoising and Complementation Framework for Multimodal Sentiment Analysis with Missing and Noisy Modalities

Yan Zhuang, Minhao Liu, Yanru Zhang, Jiawen Deng, Fuji Ren

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05735 2025-11-14 cs.AI 79%

A Comprehensive Survey on Multi-modal Conversational Emotion Recognition with Deep Learning

Yuntao Shou, Tao Meng, Wei Ai, Fangze Fu, Nan Yin, Keqin Li

机构 * Central South University of Forestry and Technology(中部林业科技大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) State University of New York(纽约州立大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

Comments 36 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09448 2025-11-13 cs.MM cs.LG 79%

MCAD: Multimodal Context-Aware Audio Description Generation For Soccer

Lipisha Chaudhary, Trisha Mittal, Subhadra Gopalakrishnan, Ifeoma Nwogu, Jaclyn Pytlarz

机构 * University at Buffalo, SUNY(布法罗大学,SUNY) Dolby Laboratories Inc.(杜比实验室公司)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22229 2025-11-13 cs.SD eess.AS 79%

Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device

Zixuan Li, Xueliang Zhang, Lei Miao, Zhipeng Yan, Ying Sun, Chong Zhu

机构 * College of Computer Science, Inner Mongolia University, China(内蒙古大学计算机科学学院) Lenovo, China(联想公司)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13979 2025-11-12 cs.CL 79%

Mixed Signals: Understanding Model Disagreement in Multimodal Empathy Detection

Maya Srikanth, Run Chen, Julia Hirschberg

机构 * Columbia University(哥伦比亚大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments To appear in Findings of IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00065 2025-11-11 cs.LG cs.AI 79%

Aligning Brain Signals with Multimodal Speech and Vision Embeddings

Kateryna Shapovalenko, Quentin Auster

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06441 2025-11-11 cs.CL cs.LG 79%

Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models

Mayank Saini, Arit Kumar Bishwas

机构 * PwC US(普华永道美国)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05497 2025-11-11 cs.IR cs.LG cs.MM 79%

Socially Aware Music Recommendation: A Multi-Modal Graph Neural Networks for Collaborative Music Consumption and Community-Based Engagement

Kajwan Ziaoddini

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05432 2025-11-10 cs.CV 79%

Shared Latent Representation for Joint Text-to-Audio-Visual Synthesis

Dogucan Yaman, Seymanur Akti, Fevziye Irem Eyiokur, Alexander Waibel

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25760 2025-11-04 cs.CV 79%

Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks

Xu Zheng, Zihao Dongfang, Lutao Jiang, Boyuan Zheng, Yulong Guo, Zhenquan Zhang, Giuliano Albanese, Runyi Yang, Mengjiao Ma, Zixin Zhang, Chenfei Liao, Dingcheng Zhen, Yuanhuiyi Lyu, Yuqian Fu, Bin Ren, Linfeng Zhang, Danda Pani Paudel, Nicu Sebe, Luc Van Gool, Xuming Hu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.09221 2025-11-03 cs.CV cs.CL cs.MM cs.SD eess.AS 79%

Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?

Yiwen Guan, Viet Anh Trinh, Vivek Voleti, Jacob Whitehill

专题命中 音频语音多模态 :multi-modal(title);分类 cs.CV、cs.CL、cs.MM

Journal ref IEEE International Conference on Multimedia and Expo (ICME), Nantes, France, 2025, pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24247 2025-10-29 cs.CL 79%

Abjad AI at NADI 2025: CATT-Whisper: Multimodal Diacritic Restoration Using Text and Speech Representations

Ahmad Ghannam, Naif Alharthi, Faris Alasmary, Kholood Al Tabash, Shouq Sadah, Lahouari Ghouti

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00937 2025-10-23 cs.DC cs.AI 79%

ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving

Haoran Qiu, Anish Biswas, Zihan Zhao, Jayashree Mohan, Alind Khare, Esha Choukse, Íñigo Goiri, Zeyu Zhang, Haiying Shen, Chetan Bansal, Ramachandran Ramjee, Rodrigo Fonseca

机构 * Microsoft Azure Research(微软Azure研究) Microsoft Research India(微软印度研究) University of Virginia(弗吉尼亚大学) Microsoft M365 Research(微软M365研究)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments Published at ACM SoCC 2025; 14 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16437 2025-10-21 eess.AS 79%

Audio-Visual Speech Enhancement for Spatial Audio - Spatial-VisualVoice and the MAVE Database

Danielle Yaffe, Ferdinand Campe, Prachi Sharma, Dorothea Kolossa, Boaz Rafaely

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15685 2025-10-20 cs.CL 79%

Leveraging LLMs for Context-Aware Implicit Textual and Multimodal Hate Speech Detection

Joshua Wolfe Brook, Ilia Markov

机构 * Computational Linguistics \& Text Mining Lab Vrije Universiteit Amsterdam

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 8 pages, 9 figures, submitted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13308 2025-10-16 eess.AS 79%

Towards Multimodal Query-Based Spatial Audio Source Extraction

Chenxin Yu, Hao Ma, Xu Li, Xiao-Lei Zhang, Mingjie Shao, Chi Zhang, Xuelong Li

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12851 2025-10-16 cs.SD cs.LG eess.AS 79%

Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models

Tsung-En Lin, Kuan-Yi Lee, Hung-Yi Lee

机构 * National Taiwan University(国立台湾大学) ASUS Open Cloud Infrastructure Software Center(ASUS开放云基础设施软件中心)

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 eess.AS

Comments Note: This preprint is a version of the paper submitted to ICASSP 2026. The author list here includes contributors who provided additional supervision and guidance. The official ICASSP submission may differ slightly in author composition

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06689 2025-10-15 cs.SD eess.AS 79%

A Fast and Lightweight Model for Causal Audio-Visual Speech Separation

Wendi Sang, Kai Li, Runxuan Yang, Jianqiang Huang, Xiaolin Hu

机构 * School of Computer Technology and Application(计算机技术与应用学院) Intelligent Computing and Application Laboratory of Qinghai Province(青海省智能计算与应用实验室) Department of Computer Science and Technology(计算机科学与技术系) Institute for AI(人工智能研究院) Tsinghua Laboratory of Brain and Intelligence (THBI)(清华大学脑与智能实验室) IDG/McGovern Institute for Brain Research(IDG/麦克戈维脑研究学院) Tsinghua University(清华大学) Chinese Institute for Brain Research (CIBR)(中国脑科学研究院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted by ECAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05417 2025-10-08 cs.HC cs.AI 79%

Exploring Student Choice and the Use of Multimodal Generative AI in Programming Learning

Xinying Hou, Ruiwei Xiao, Runlong Ye, Michael Liut, John Stamper

机构 * University of Michigan(密歇根大学) Carnegie Mellon University(卡内基梅隆大学) University of Toronto(多伦多大学) University of Toronto Mississauga(多伦多大学滑铁卢分校)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments 7 pages, accepted to SIGCSE2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07748 2025-10-07 eess.AS 79%

Leveraging Self-Supervised Audio-Visual Pretrained Models to Improve Vocoded Speech Intelligibility in Cochlear Implant Simulation

Richard Lee Lai, Jen-Cheng Hou, I-Chun Chern, Kuo-Hsuan Hung, Yi-Ting Chen, Mandar Gogate, Tughrul Arslan, Amir Hussain, Yu Tsao

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13624 2025-10-01 cs.SD eess.AS 79%

Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement

Rong Chao, Wenze Ren, You-Jin Li, Kuo-Hsuan Hung, Sung-Feng Huang, Szu-Wei Fu, Wen-Huang Cheng, Yu Tsao

机构 * Academia Sinica(台湾“中央研究院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted to Interspeech 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22698 2025-09-30 cs.RO cs.AI 79%

Advancing Audio-Visual Navigation Through Multi-Agent Collaboration in 3D Environments

Hailong Zhang, Yinfeng Yu, Liejun Wang, Fuchun Sun, Wendong Zheng

机构 * Xinjiang Multimodal Intelligent Processing and Information Security Engineering Technology Research Center(新疆多模态智能处理与信息安全工程技术研究中心) School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) School of Electrical Engineering and Automation, Tianjin University of Technology(天津理工大学电气工程与自动化学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI

Comments Main paper (15 pages). Accepted for publication by ICONIP( International Conference on Neural Information Processing) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21749 2025-09-29 cs.CL cs.SD 79%

Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models

Zhen Xiong, Yujun Cai, Zhecheng Li, Junsong Yuan, Yiwei Wang

机构 * University of Southern California(南加州大学) University of Queensland(昆士兰大学) University of California, San Diego(加州大学圣地亚哥分校) University of Buffalo(布法罗大学) University of California, Merced(加州大学默塞德分校)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16378 2025-09-29 cs.CY cs.CL 79%

Longitudinal and Multimodal Recording System to Capture Real-World Patient-Clinician Conversations for AI and Encounter Research: Protocol

Misk Al Zahidy, Kerly Guevara Maldonado, Luis Vilatuna Andrango, Ana Cristina Proano, Ana Gabriela Claros, Maria Lizarazo Jimenez, David Toro-Tobon, Victor M. Montori, Oscar J. Ponce-Ponte, Juan P. Brito

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 23 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏