arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-22 至 2025-09-22 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 6 篇

2509.16193 2025-09-22 eess.AS 89%

Are Multimodal Foundation Models All That Is Needed for Emofake Detection?

Mohd Mujtaba Akhtar, Girish, Orchid Chetia Phukan, Swarup Ranjan Behera, Pailla Balakrishna Reddy, Ananda Chandra Nayak, Sanjib Kumar Nayak, Arun Balaji Buduru

专题命中 音频语音多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);cross-modal(abstract);分类 eess.AS

Comments Accepted to APSIPA-ASC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16025 2025-09-22 cs.CL cs.AI 88%

Session-Level Spoken Language Assessment with a Multimodal Foundation Model via Multi-Target Learning

Hong-Yun Lin, Jhen-Ke Lin, Chung-Chun Wang, Hao-Chien Lu, Berlin Chen

机构 * Department of Computer Science(计算机科学系) Information Engineering, National Taiwan Normal University(信息工程,台湾正常大学)

专题命中 音频语音多模态 :multimodal(title,abstract);multimodal foundation model(title,abstract);分类 cs.CL、cs.AI

Comments Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15476 2025-09-22 cs.CL cs.MM 86%

Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding

Zhu Li, Xiyuan Gao, Yuqing Zhang, Shekhar Nayak, Matt Coler

机构 * University of Groningen, The Netherlands(Groningen大学,荷兰)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15661 2025-09-22 cs.SD cs.AI cs.CL eess.AS 85%

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

Qiaolin Wang, Xilin Jiang, Linyang He, Junkai Wu, Nima Mesgarani

机构 * Columbia University(哥伦比亚大学) University of Washington(华盛顿大学)

专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16023 2025-09-22 eess.AS 79%

Interpreting the Role of Visemes in Audio-Visual Speech Recognition

Aristeidis Papadopoulos, Naomi Harte

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted into Automatic Speech Recognition and Understanding- ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15775 2025-09-22 cs.SD eess.AS 70%

EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model

Yiqing Yang, Man-Wai Mak

机构 * Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(电子与电气工程系,香港理工大学)

专题命中 音频语音多模态 :multimodal(abstract);MLLM(abstract);分类 eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏