arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-09 至 2025-09-09 共收录 4 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4 篇

2402.12226 2025-09-09 cs.CL cs.AI cs.CV cs.LG 85%

AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang, Xipeng Qiu

机构 * Fudan University(复旦大学) Multimodal Art Projection Research Community(多模态艺术投影研究社区) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 音频语音多模态 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 28 pages, 16 figures, under review, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06074 2025-09-09 cs.CL 79%

Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis

Zhenqi Jia, Rui Liu, Berrak Sisman, Haizhou Li

机构 * Inner Mongolia University(内蒙古大学) Center for Language and Speech Processing (CLSP)(语言与语音处理中心) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)人工智能学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06598 2025-09-09 eess.AS cs.AI cs.LG eess.IV eess.SP 79%

Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos

Davide Berghi, Philip J. B. Jackson

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.AI、eess.AS

Comments arXiv admin note: substantial text overlap with arXiv:2507.04845

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06382 2025-09-09 cs.HC 78%

Context-Adaptive Hearing Aid Fitting Advisor through Multi-turn Multimodal LLM Conversation

Yingke Ding, Zeyu Wang, Xiyuxing Zhang, Hongbin Chen, Zhenan Xu

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Ubicomp Companion 2025

详情

展开后加载摘要…

URL PDF HTML 收藏