arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-12 至 2025-08-12 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 7 篇

2508.07608 2025-08-12 cs.MM cs.CV cs.SD eess.AS 85%

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition

Junxiao Xue, Xiaozhen Liu, Xuecheng Wu, Xinyi Yin, Danlei Huang, Fei Yu

机构 * Zhengzhou University(郑州大学) Zhejiang Lab(浙江实验室) Xi'an Jiaotong University(西安交通大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted by the ACM MM 2025 Workshop on SVC

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06902 2025-08-12 cs.CV 83%

eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos

Xuecheng Wu, Dingkang Yang, Danlei Huang, Xinyi Yin, Yifan Wang, Jia Zhang, Jiayu Nie, Liangyu Fu, Yang Liu, Junxiao Xue, Hadi Amirpour, Wei Zhou

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(计算机科学与技术学院,西安交通大学) College of Intelligent Robotics and Advanced Manufacturing, Fudan University & ByteDance(智能机器人与先进制造学院,复旦大学 & 字节跳动) School of Cyber Science and Engineering, Zhengzhou University(网络科学与工程学院,郑州大学) Institute of Advanced Technology, University of Science and Technology of China(先进技术研究院,中国科学技术大学) Inspur Electronic Information Industry Co., Ltd(Inspur电子信息产业有限公司) Department of Computer Science, The University of Toronto(计算机科学系,多伦多大学) Research Center for Space Computing System, Zhejiang Lab(空间计算系统研究中心,浙江实验室) Institute of Information Technology, University of Klagenfurt(信息技术学院,克雷格福特大学) School of Computer Science and Informatics, Cardiff University(计算机科学与信息学院,卡迪夫大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07337 2025-08-12 eess.AS cs.CV 81%

KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features

Ivan Kukanov, Jun Wah Ng

机构 * KLASS Engineering and Solutions Singapore(KLASS工程与解决方案新加坡)

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 cs.CV、eess.AS

Comments 7 pages, accepted to the 33rd ACM International Conference on Multimedia (MM'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08042 2025-08-12 cs.IR cs.AI 79%

Multi-modal Adaptive Mixture of Experts for Cold-start Recommendation

Van-Khang Nguyen, Duc-Hoang Pham, Huy-Son Nguyen, Cam-Van Thi Nguyen, Hoang-Quynh Le, Duc-Trong Le

机构 * VNU University of Engineering and Technology(越南工程大学) Delft University of Technology(代尔夫特理工大学)

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02516 2025-08-12 cs.CV 79%

Engagement Prediction of Short Videos with Large Multimodal Models

Wei Sun, Linhan Cao, Yuqin Cao, Weixia Zhang, Wen Wen, Kaiwei Zhang, Zijian Chen, Fangfang Lu, Xiongkuo Min, Guangtao Zhai

机构 * East China Normal University(华东师范大学) Shanghai Jiao Tong University(上海交通大学) City University of Hong Kong(香港城市大学) Shanghai University of Electric Power(上海电力大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments The proposed method achieves first place in the ICCV VQualA 2025 EVQA-SnapUGC Challenge on short-form video engagement prediction

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07625 2025-08-12 cs.CV 74%

A Trustworthy Method for Multimodal Emotion Recognition

Junxiao Xue, Xiaozhen Liu, Jie Wang, Xuecheng Wu, Bin Wu

机构 * Big Data Mining and Analytics, xxxxxxx 20xx, x(x): xxx-xxx(大数据挖掘与分析,xxxxx 20xx, x(x): xxx-xxx)

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV

Comments Accepted for publication in Big Data Mining and Analytics (BDMA), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07973 2025-08-12 cs.SD cs.CL eess.AS 62%

Joint Transcription of Acoustic Guitar Strumming Directions and Chords

Sebastian Murgul, Johannes Schimper, Michael Heizmann

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted to the 26th International Society for Music Information Retrieval Conference (ISMIR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏