arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-19 至 2025-12-19 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 5 篇

2512.01185 2025-12-19 cs.CR 85%

DefenSee: Dissecting Threat from Sight and Text -- A Multi-View Defensive Pipeline for Multi-modal Jailbreaks

DefenSee:从视觉和文本中解构威胁——一种多视图防御管道用于多模态对抗突破

Zihao Wang, Kar Wai Fok, Vrizlynn L. L. Thing

专题命中 音频语音多模态 :multi-modal(title,abstract);MLLM(abstract);cross-modal(abstract)

AI总结 DefenSee通过图像变体转录和跨模态一致性检查,提供一种多模态防御方法,有效降低多模态对抗攻击的成功率,提升模型鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16250 2025-12-19 cs.AI cs.MA 83%

AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding

AMUSE:面向代理多说话者理解的音频-视觉基准与对齐框架

Sanjoy Chowdhury, Karren D. Yang, Xudong Liu, Fartash Faghri, Pavan Kumar Anasosalu Vasu, Oncel Tuzel, Dinesh Manocha, Chun-Liang Li, Raviteja Vemulapalli

机构 * University of Maryland, College Park(马里兰大学学院公园分校) Apple(苹果公司)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 AMUSE提出一个面向多说话人理解的音频-视觉基准与对齐框架RAFT,通过代理推理提升多模态模型能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14491 2025-12-19 cs.LG cs.MM 79%

Multimodal Methods for Analyzing Learning and Training Environments: A Systematic Literature Review

多模态方法用于分析学习与训练环境:系统文献综述

Clayton Cohn, Eduardo Davalos, Caleb Vatral, Joyce Horn Fonteles, Hanchen David Wang, Austin Coursey, Surya Rayala, Ashwin T S, Meiyi Ma, Gautam Biswas

机构 * Vanderbilt University(范德比大学) Tennessee State University(田纳西州立大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

AI总结 本文综述了多模态方法在学习与训练环境中的应用,提出分类法和框架,揭示了多模态整合对行为分析的价值及现存挑战。

Comments Submitted to ACM Computing Surveys. Currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.07741 2025-12-19 cs.LG cs.SD 78%

A multimodal Bayesian Network for symptom-level depression and anxiety prediction from voice and speech data

一种多模态贝叶斯网络用于从语音和语音数据中预测症状层面的抑郁和焦虑

Agnes Norbury, George Fairs, Alexandra L. Georgescu, Matthew M. Nour, Emilia Molimpakis, Stefano Goria

机构 * thymia Limited(thymia有限公司) Institute of Psychiatry, Psychology & Neuroscience, King’s College London(心理学与神经科学研究院,伦敦国王学院) Department of Psychiatry, University of Oxford(牛津大学精神病学系) Max Planck UCL Centre for Computational Psychiatry and Ageing, University College London(Max Planck大学学院计算精神病学与衰老中心,伦敦大学学院)

专题命中 音频语音多模态 :multimodal(title,abstract)

AI总结 本文提出了一种多模态贝叶斯网络模型,用于从语音和语音数据中预测抑郁和焦虑症状,通过评估模型性能和公平性,展示了其在临床应用中的潜力。

Journal ref Scientific Reports (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12424 2025-12-19 cs.LG cs.AI cs.IR 70%

Multi-Modality Collaborative Learning for Sentiment Analysis

多模态协作学习用于情感分析

Shanmin Wang, Chengguang Liu, Qingshan Liu

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出多模态协作学习框架,通过解耦模块和策略模型提升跨模态情感特征学习,实验证明在四个数据库上性能显著提升。

Comments The method has flaws, especially with the decoupling module. During the decoupling process, the heterogeneity of the three modal data and the differences in distribution were not taken into account

详情

展开后加载摘要…

URL PDF HTML 收藏