arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-22 至 2025-08-22 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 6 篇

2508.12227 2025-08-22 cs.CL 79%

Arabic Multimodal Machine Learning: Datasets, Applications, Approaches, and Challenges

Abdelhamid Haouhat, Slimane Bellaouar, Attia Nehar, Hadda Cherroun, Ahmed Abdelali

机构 * Ziane Achour University(赞赞·阿赫尔大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15565 2025-08-22 cs.SD 78%

Any-to-any Speaker Attribute Perturbation for Asynchronous Voice Anonymization

Liping Chen, Chenyang Guo, Rui Wang, Kong Aik Lee, Zhenhua Ling

机构 * University of Science and Technology of China(中国科学技术大学) Hong Kong Polytechnic University(香港理工大学)

专题命中 音频语音多模态 :any-to-any(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14976 2025-08-22 cs.LG 78%

Aura-CAPTCHA: A Reinforcement Learning and GAN-Enhanced Multi-Modal CAPTCHA System

Joydeep Chandra, Prabal Manhas, Ramanjot Kaur, Rashi Sahay

机构 * Department of Computer Science and Engineering, Chandigarh University, Mohali, Punjab, India(昌迪加尔大学计算机科学与工程系,莫哈利,旁遮普,印度) Department of Computer Science and Engineering, Manav Rachna International Institute of Research and Studies(曼纳瓦拉国际研究与学习研究所计算机科学与工程系)

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14912 2025-08-22 cs.IR 78%

Multimodal Recommendation via Self-Corrective Preference Alignmen

Yalong Guan, Xiang Chen, Mingyang Wang, Xiangyu Wu, Lihao Liu, Chao Qi, Shuang Yang, Tingting Gao, Guorui Zhou, Changjian Chen

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15407 2025-08-22 cs.CL cs.AI 73%

When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models

Cheng Wang, Gelei Deng, Xianglin Yang, Han Qiu, Tianwei Zhang

机构 * National University of Singapore(国立新加坡大学) Nanyang Technological University(南洋理工大学) Tsinghua University(清华大学)

专题命中 音频语音多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12918 2025-08-22 cs.SD 50%

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation

Lei Zhao, Rujin Chen, Chi Zhang, Xiao-Lei Zhang, Xuelong Li

机构 * School of Marine Science and Technology, Northwestern Polytechnical University(海洋科学与技术学院,西北工业大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) Research and Development Institute of Northwestern Polytechnical University in Shenzhen, China(西北工业大学深圳研发院,中国)

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏