arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46352 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4587 篇

2504.19611 2026-02-10 cs.HC 78%

Scene2Hap: Generating Scene-Wide Haptics for VR from Scene Context with Multimodal LLMs

Scene2Hap: 从场景上下文生成VR场景中的全域触觉反馈

Arata Jingu, Easa AliAbbasi, Sara Safaee, Paul Strohmeier, Jürgen Steimle

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 Scene2Hap通过多模态大语言模型生成VR场景中的触觉反馈,提升沉浸感与真实感。

Comments Accepted at CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17638 2026-01-27 cs.CR 78%

FOCA: Multimodal Malware Classification via Hyperbolic Cross-Attention

FOCA:基于双模态的恶意软件分类方法

Nitin Choudhury, Bikrant Bikram Pratap Maurya, Orchid Chetia Phukan, Arun Balaji Buduru

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 FOCA通过双曲空间中的交叉注意力机制和双曲投影模块,实现音频和视觉模态的恶意软件分类,展现出优越的分类性能。

Comments Accepted to the International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14227 2026-01-21 cs.SD 78%

Transformer Architectures for Respiratory Sound Analysis and Multimodal Diagnosis

用于呼吸声分析和多模态诊断的Transformer架构

Theodore Aptekarev, Vladimir Sokolovsky, Gregory Furman

机构 * Ben Gurion University of the Negev(本· Gurion 农业大学)

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 本文提出基于Transformer的AST和VLM模型,用于呼吸声分析和多模态诊断,AST在哮喘检测中达到97%准确率,VLM整合临床信息提升诊断能力。

Comments 7 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01166 2026-01-19 cs.SD 78%

Hearing More with Less: Multi-Modal Retrieval-and-Selection Augmented Conversational LLM-Based ASR

听更清晰,用更少:多模态检索与选择增强的对话式LLM基于ASR

Bingshen Mu, Hexin Liu, Hongfei Xue, Kun Wei, Lei Xie

专题命中 音频语音多模态 :multi-modal(title,abstract)

AI总结 本文提出MARS方法,通过多模态检索与选择提升对话式LLM-ASR的准确性,仅用1.5K小时数据即可超越训练于179K小时数据的顶级系统。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07808 2026-01-15 cs.RO 78%

AcoustoBots: A swarm of robots for acoustophoretic multimodal interactions

AcoustoBots:用于声学电势多模交互的机器人群

Narsimlu Kemsaram, James Hardwick, Jincheng Wang, Bonot Gautam, Ceylan Besevli, Giorgos Christopoulos, Sourabh Dogra, Lei Gao, Akin Delibasi, Diego Martinez Plasencia, Orestis Georgiou, Marianna Obrist, Ryuji Hirayama, Sriram Subramanian

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 AcoustoBots通过可移动相位阵列和BeadDispenserBot实现声学电势多模交互,提升机器人群的灵活性和交互能力。

Journal ref Frontiers in Robotics and AI, 12:1537101, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07741 2025-12-19 cs.LG cs.SD 78%

A multimodal Bayesian Network for symptom-level depression and anxiety prediction from voice and speech data

一种多模态贝叶斯网络用于从语音和语音数据中预测症状层面的抑郁和焦虑

Agnes Norbury, George Fairs, Alexandra L. Georgescu, Matthew M. Nour, Emilia Molimpakis, Stefano Goria

机构 * thymia Limited(thymia有限公司) Institute of Psychiatry, Psychology & Neuroscience, King’s College London(心理学与神经科学研究院,伦敦国王学院) Department of Psychiatry, University of Oxford(牛津大学精神病学系) Max Planck UCL Centre for Computational Psychiatry and Ageing, University College London(Max Planck大学学院计算精神病学与衰老中心,伦敦大学学院)

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 本文提出了一种多模态贝叶斯网络模型,用于从语音和语音数据中预测抑郁和焦虑症状,通过评估模型性能和公平性,展示了其在临床应用中的潜力。

Journal ref Scientific Reports (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21429 2025-12-12 cs.LG 78%

Deception Detection in Dyadic Exchanges Using Multimodal Machine Learning: A Study on a Swedish Cohort

利用多模态机器学习检测双人交流中的欺骗:一项针对瑞典群体的研究

Thomas Jack Samuels, Franco Rugolon, Stephan Hau, Lennart Högman

机构 * Department of Psychology(心理学系) Department of Computer and Systems Sciences(计算机与系统科学系) Stockholm University(斯德哥尔摩大学)

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 本研究利用多模态机器学习分析双人互动中的欺骗检测,发现结合语音和面部信息及双方数据可提高检测准确率至71%。

Comments 40 pages, 2 figures, 2 tables. To be submitted in Behavior Research Methods

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09504 2025-12-11 cs.SD 78%

DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance

DMP-TTS: 解耦多模态提示的可控文本到语音系统 with 链式引导

Kang Yin, Chunyu Qiang, Sirui Zhao, Xiaopeng Wang, Yuzhe Liang, Pengfei Cai, Tong Xu, Chen Zhang, Enhong Chen

机构 * University of Science and Technology of China(科学技术大学)

专题命中 音频语音多模态 :multi-modal(title,abstract)

AI总结 DMP-TTS通过解耦多模态提示和链式引导技术,实现了更强的文本到语音风格可控性,同时保持良好的可懂度和自然度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04275 2025-12-01 cs.DC 78%

DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models

DistTrain: 通过解耦训练解决多模态大语言模型中的模型与数据异质性

Zili Zhang, Yinmin Zhong, Yimin Jiang, Hanpeng Hu, Jianjian Sun, Zheng Ge, Yibo Zhu, Daxin Jiang, Xin Jin

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 DistTrain通过解耦模型编排和数据预处理,提升多模态大语言模型的训练效率和资源利用率。

Comments SIGCOMM 2025 (https://dl.acm.org/doi/10.1145/3718958.3750472)

Journal ref SIGCOMM'25: Proceedings of the ACM SIGCOMM 2025 Conference Pages 24-38

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16225 2025-11-21 cs.LG 78%

Real-Time Inference for Distributed Multimodal Systems under Communication Delay Uncertainty

在通信延迟不确定性下的分布式多模态系统实时推理

Victor Croisfelt, João Henrique Inacio de Souza, Shashi Raj Pandey, Beatriz Soret, Petar Popovski

机构 * SNS JU project 6G-GOALS under the EU’s Horizon Europe program(欧盟地平线欧洲计划下的SNS JU项目6G-GOALS)

专题命中 音频语音多模态 :multimodal(title);audio-visual(abstract)

AI总结 本文提出了一种基于自适应时间窗口整合的非阻塞推理方法,以应对通信延迟不确定性,提升分布式多模态系统的实时推理性能。

Comments 6 pages, 3 figures, submitted to IEEE ICC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10011 2025-11-18 cs.CY 78%

Reinforcing Trustworthiness in Multimodal Emotional Support Systems

Huy M. Le, Dat Tien Nguyen, Ngan T. T. Vo, Tuan D. Q. Nguyen, Nguyen Binh Le, Duy Minh Ho Nguyen, Daniel Sonntag, Lizi Liao, Binh T. Nguyen

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09989 2025-11-14 cs.LG 78%

Towards Robust Multimodal Learning in the Open World

Fushuo Huo

机构 * Department of Computing(计算系)

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09039 2025-11-13 cs.LG cs.CY cs.HC 78%

Fairness-Aware Few-Shot Learning for Audio-Visual Stress Detection

Anushka Sanjay Shelke, Aditya Sneh, Arya Adyasha, Haroon R. Lone

专题命中 音频语音多模态 :audio-visual(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05304 2025-11-10 cs.HC 78%

psiUnity: A Platform for Multimodal Data-Driven XR

Akhil Ajikumar, Sahil Mayenkar, Steven Yoo, Sakib Reza, Mohsen Moghaddam

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26830 2025-11-03 cs.LG cs.CR 78%

SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation

Guangzhi Su, Shuchang Huang, Yutong Ke, Zhuohang Liu, Long Qian, Kaizhu Huang

机构 * Independent Researcher(独立研究者)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10396 2025-10-20 cs.SD 78%

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations

Wenxiang Guo, Changhao Pan, Zhiyuan Zhu, Xintong Hu, Yu Zhang, Li Tang, Rui Yang, Han Wang, Zongbao Zhang, Yuhan Wang, Yixuan Chen, Hankun Xu, Ke Xu, Pengfei Fan, Zhetao Chen, Yanhao Yu, Qiange Huang, Fei Wu, Zhou Zhao

机构 * Zhejiang University(浙江大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11596 2025-10-14 cs.HC 78%

GlobalizeEd: A Multimodal Translation System that Preserves Speaker Identity in Academic Lectures

Hoang-Son Vo, Karina Kolmogortseva, Ngumimi Karen Iyortsuun, Hong-Duyen Vo, Soo-Hyung Kim

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22425 2025-10-13 cs.SD 78%

From Coarse to Fine: Recursive Audio-Visual Semantic Enhancement for Speech Separation

Ke Xue, Rongfei Fan, Lixin, Dawei Zhao, Chao Zhu, Han Hu

机构 * School of Cyberspace Science and Technology(网络空间科学与技术学院) Beijing Institute of Technology(北京理工大学) Sun Yat-sen University(中山大学) Qilu University of Technology(齐鲁工业大学) Shandong Computer Science Center(山东计算机科学中心) School of Information and Electronics(信息电子学院)

专题命中 音频语音多模态 :audio-visual(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07839 2025-10-09 cs.RO 78%

Touch Speaks, Sound Feels: A Multimodal Approach to Affective and Social Touch from Robots to Humans

Qiaoqiao Ren, Tony Belpaeme

机构 * Faculty of Engineering and Architecture, IDLab-AIRO, Ghent University – imec, Technologiepark 126, 9052 Gent, Belgium(工程与建筑学院,IDLab-AIRO,根特大学–imec,Technologiepark 126,9052 Gent,比利时)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06872 2025-10-09 cs.HC 78%

Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-Oz

Frederic Gmeiner, Kenneth Holstein, Nikolas Martelaro

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 18 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19600 2025-09-25 cs.HC 78%

vashTimer: A Multi-Purpose, Multimodal Mobile App For Maintaining Passage of Time by means of Visual, Auditory, Speech, and Haptic Alerts

Aziz N Zeidieh, Sanchita S. Kamath, JooYoung Seo

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18706 2025-09-24 cs.HC 78%

M4SER: Multimodal, Multirepresentation, Multitask, and Multistrategy Learning for Speech Emotion Recognition

Jiajun He, Xiaohan Shi, Cheng-Hung Hu, Jinyi Mi, Xingfeng Li, Tomoki Toda

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted by IEEE Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16920 2025-09-23 cs.RO cs.HC 78%

SwarmChat: An LLM-Based, Context-Aware Multimodal Interaction System for Robotic Swarms

Ettilla Mohiuddin Eumi, Hussein Abbass, Nadine Marcus

机构 * School of Systems & Computing, UNSW Canberra(系统与计算学院,UNSW堪培拉) School of Computer Science and Engineering, UNSW Sydney(计算机科学与工程学院,UNSW悉尼)

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This paper has been accepted and presented at the 16th International Conference on Swarm Intelligence (ICSI 2025), held on July 11-15, 2025, in Yokohama, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02359 2025-09-12 cs.LG 78%

Attribution Regularization for Multimodal Paradigms

Sahiti Yerramilli, Jayant Sravan Tamarapalli, Jonathan Francis, Eric Nyberg

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08689 2025-09-11 cs.HC 78%

Augmenting speech transcripts of VR recordings with gaze, pointing, and visual context for multimodal coreference resolution

Riccardo Bovo, Frederik Brudy, George Fitzmaurice, Fraser Anderson

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06382 2025-09-09 cs.HC 78%

Context-Adaptive Hearing Aid Fitting Advisor through Multi-turn Multimodal LLM Conversation

Yingke Ding, Zeyu Wang, Xiyuxing Zhang, Hongbin Chen, Zhenan Xu

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Ubicomp Companion 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15632 2025-08-26 cs.SD 78%

ASCMamba: Multimodal Time-Frequency Mamba for Acoustic Scene Classification

Bochao Sun, Dong Wang, ZhanLong Yang, Jun Yang, Han Yin

机构 * School of Marine Science and Technology, Northwestern Polytechnical University, Xi’an, China(海洋科学与技术学院,西北工业大学,西安,中国) School of Automation, Northwestern Polytechnical University, Xi’an, China(自动化学院,西北工业大学,西安,中国) School of Electrical Engineering, KAIST, Daejeon, Republic of Korea(电气工程学院,韩国成均馆大学,大田,韩国)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16320 2025-08-25 physics.ed-ph 78%

AI-Supported Mini-Labs: Combining Smartphone-Based Experiments and Multimodal AI

Jochen Kuhn, David J. Rakestraw, Stefan Küchemann, Patrik Vogt

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15565 2025-08-22 cs.SD 78%

Any-to-any Speaker Attribute Perturbation for Asynchronous Voice Anonymization

Liping Chen, Chenyang Guo, Rui Wang, Kong Aik Lee, Zhenhua Ling

机构 * University of Science and Technology of China(中国科学技术大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 音频语音多模态 :any-to-any(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14976 2025-08-22 cs.LG 78%

Aura-CAPTCHA: A Reinforcement Learning and GAN-Enhanced Multi-Modal CAPTCHA System

Joydeep Chandra, Prabal Manhas, Ramanjot Kaur, Rashi Sahay

机构 * Department of Computer Science and Engineering, Chandigarh University, Mohali, Punjab, India(昌迪加尔大学计算机科学与工程系,莫哈利,旁遮普,印度) Department of Computer Science and Engineering, Manav Rachna International Institute of Research and Studies(曼纳瓦拉国际研究与学习研究所计算机科学与工程系)

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏