arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-16 至 2025-10-16 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 9 篇

2507.09945 2025-10-16 cs.MM cs.CV 86%

ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization

Huilai Li, Yonghao Dang, Ying Xing, Yiming Wang, Jianqin Yin

机构 * School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications(智能工程与自动化学院,北京邮电大学) School of Artificial Intelligence, Beijing University of Posts and Telecommunications(人工智能学院,北京邮电大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13182 2025-10-16 cs.LG 82%

Information-Theoretic Criteria for Knowledge Distillation in Multimodal Learning

Rongrong Xie, Yizhou Xu, Guido Sanguinetti

机构 * Scuola Internazionale Superiore di Studi Avanzati (SISSA)(国际先进研究高等学院) École Polytechnique Fédérale de Lausanne (EPFL)(日内瓦联邦理工学院)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13281 2025-10-16 eess.AS cs.CL cs.LG 81%

Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses

Sungnyun Kim, Kangwook Jang, Sungwoo Cho, Joon Son Chung, Hoirin Kim, Se-Young Yun

机构 * Kim Jaechul Graduate School of AI, KAIST(金 Jaechul人工智能研究生院,韩国科学技术院) School of Electrical Engineering, KAIST(电气工程学院,韩国科学技术院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CL、eess.AS

Comments Preprint work

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13308 2025-10-16 eess.AS 79%

Towards Multimodal Query-Based Spatial Audio Source Extraction

Chenxin Yu, Hao Ma, Xu Li, Xiao-Lei Zhang, Mingjie Shao, Chi Zhang, Xuelong Li

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12851 2025-10-16 cs.SD cs.LG eess.AS 79%

Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models

Tsung-En Lin, Kuan-Yi Lee, Hung-Yi Lee

机构 * National Taiwan University(国立台湾大学) ASUS Open Cloud Infrastructure Software Center(ASUS开放云基础设施软件中心)

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 eess.AS

Comments Note: This preprint is a version of the paper submitted to ICASSP 2026. The author list here includes contributors who provided additional supervision and guidance. The official ICASSP submission may differ slightly in author composition

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13351 2025-10-16 cs.CL cs.AI 62%

Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems

Karthik Avinash, Nikhil Pareek, Rishav Hada

机构 * FutureAGI Inc.(未来人工智能公司)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13344 2025-10-16 cs.SD cs.CL 57%

UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE

Zhenyu Liu, Yunxin Li, Xuanyu Zhang, Qixun Teng, Shenyuan Jiang, Xinyu Chen, Haoyuan Shi, Jinchao Li, Qi Wang, Haolan Chen, Fanbo Meng, Mingjun Zhao, Yu Xu, Yancheng He, Baotian Hu, Min Zhang

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05803 2025-10-16 cs.GR cs.CV 57%

PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis

Yihuan Huang, Jiajun Liu, Yanzhen Ren, Jun Xue, Wuyang Liu, Zongkun Sun

机构 * Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航空航天信息安全部分和可信计算重点实验室、教育部、网络安全科学与工程学院、武汉大学) School of Cyber Science and Engineering, Wuhan University(网络安全科学与工程学院、武汉大学) School of Police Information, Shandong Police College(警务信息学院、山东警察学院)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13558 2025-10-16 cs.SD 50%

Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module

Ruitao Feng, Bixi Zhang, Sheng Liang, Zheng Yuan

机构 * The University of Hong Kong, Fauclty of Science, Hong Kong(香港大学科学学院) Aix-Marseille University, Laboratoire Parole et Langage (LPL), France(艾克斯-马赛大学语言与言语实验室(LPL))

专题命中 音频语音多模态 :multimodal(abstract)

Comments 5 pages, 1 figures. Code is available at: https://github.com/forfrt/SteerMoE. Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏