arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-04 至 2025-08-04 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 6 篇

2508.00632 2025-08-04 cs.AI cs.MA cs.MM 84%

Multi-Agent Game Generation and Evaluation via Audio-Visual Recordings

Alexia Jolicoeur-Martineau

机构 * Samsung SAIL Montréal(三星SAIL蒙特利尔)

专题命中 音频语音多模态 :audio-visual(title,abstract);omni-modal(abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00760 2025-08-04 cs.CL cs.AI 81%

MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations

Qiyao Xue, Yuchen Dou, Ryan Shi, Xiang Lorraine Li, Wei Gao

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00784 2025-08-04 cs.AI 79%

Unraveling Hidden Representations: A Multi-Modal Layer Analysis for Better Synthetic Content Forensics

Tom Or, Omri Azencot

机构 * Ben Gurion University of the Negev(本· Gurion 内盖夫大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00391 2025-08-04 cs.CV eess.AS 62%

Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition

Guanjie Huang, Danny H. K. Tsang, Shan Yang, Guangzhi Lei, Li Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tencent AI Lab(腾讯AI实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、eess.AS

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00160 2025-08-04 cs.HC cs.AI cs.SD eess.AS 62%

DeformTune: A Deformable XAI Music Prototype for Non-Musicians

Ziqing Xu, Nick Bryan-Kinns

机构 * Creative Computing Institute, University of the Arts London(创意计算研究所,伦敦艺术大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments In Proceedings of Explainable AI for the Arts Workshop 2025 (XAIxArts 2025) arXiv:2406.14485

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00205 2025-08-04 cs.CV 57%

Learning Personalised Human Internal Cognition from External Expressive Behaviours for Real Personality Recognition

Xiangyu Kong, Hengde Zhu, Haoqin Sun, Zhihao Guo, Jiayan Gu, Xinyi Ni, Wei Zhang, Shizhe Liu, Siyang Song

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏