arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-08 至 2025-08-08 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 5 篇

2508.00733 2025-08-08 cs.SD cs.CV cs.MM eess.AS 85%

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Le Wang, Jun Wang, Chunyu Qiang, Feng Deng, Chen Zhang, Di Zhang, Kun Gai

机构 * China University of Mining and Technology(中国矿业大学) Kuaishou Technology(快手科技)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments 12 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05087 2025-08-08 cs.MM cs.AI cs.CL cs.CR 82%

JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual Steering

Renmiao Chen, Shiyao Cui, Xuancheng Huang, Chengwei Pan, Victor Shea-Jay Huang, QingLin Zhang, Xuan Ouyang, Zhexin Zhang, Hongning Wang, Minlie Huang

机构 * CoAI group, DCST, Tsinghua University(清华大学DCST学院) Beihang University(北航大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

Comments 10 pages, 3 tables, 2 figures, to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04723 2025-08-08 cs.SD cs.AI eess.AS 73%

Wearable Music2Emotion : Assessing Emotions Induced by AI-Generated Music through Portable EEG-fNIRS Fusion

Sha Zhao, Song Yi, Yangxuan Zhou, Jiadong Pan, Jiquan Wang, Jie Xia, Shijian Li, Shurong Dong, Gang Pan

机构 * Zhejiang University(浙江大学) Hangzhou RongNao Technology Co., Ltd(杭州融脑科技有限公司)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04585 2025-08-08 eess.AS 57%

UniTalker: Conversational Speech-Visual Synthesis

Yifan Hu, Rui Liu, Yi Ren, Xiang Yin, Haizhou Li

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 15 pages, 8 figures, Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03536 2025-08-08 eess.AS 57%

Overview of Automatic Speech Analysis and Technologies for Neurodegenerative Disorders: Diagnosis and Assistive Applications

Shakeel A. Sheikh, Md. Sahidullah, Ina Kodrasi

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Published in IEEE Journal of Selected Topics in Signal Processing

Journal ref https://ieeexplore.ieee.org/abstract/document/11086511/

详情

展开后加载摘要…

URL PDF HTML 收藏