arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-14 至 2025-10-14 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 9 篇

2510.10051 2025-10-14 cs.CV 85%

Complementary and Contrastive Learning for Audio-Visual Segmentation

Sitong Gong, Yunzhi Zhuge, Lu Zhang, Pingping Zhang, Huchuan Lu

机构 * School of Information and Communication Engineering, Dalian University of Technology(信息与通信工程学院,大连理工大学) School of Future Technology and the School of Artificial Intelligence, Dalian University of Technology(未来技术学院和人工智能学院,大连理工大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01181 2025-10-14 cs.AI cs.CV cs.MM cs.SD eess.AS 83%

Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning

Zhiyuan Han, Beier Zhu, Yanlong Xu, Peipei Song, Xun Yang

机构 * University of Science and Technology of China(科学技术大学) Nanyang Technological University(南洋理工大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments ACM Multimedia 2025 Oral Code: https://github.com/ZhiyuanHan-Aaron/MoSEAR Project Page: https://zhiyuanhan-aaron.github.io/MoSEAR-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05295 2025-10-14 cs.SD cs.AI cs.MM 82%

AUREXA-SE: Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement

M. Sajid, Deepanshu Gupta, Yash Modi, Sanskriti Jain, Harshith Jai Surya Ganji, A. Rahaman, Harshvardhan Choudhary, Nasir Saleem, Amir Hussain, M. Tanveer

机构 * Indian Institute of Technology Indore(印度理工学院印第安纳分校) School of Computing, Edinburgh Napier University(爱丁堡纳皮尔大学计算机学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.AI、cs.MM

Journal ref INTERSPEECH 2025 - 4th COG-MHEAR Workshop on Audio-Visual Speech Enhancement (AVSEC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11596 2025-10-14 cs.HC 78%

GlobalizeEd: A Multimodal Translation System that Preserves Speaker Identity in Academic Lectures

Hoang-Son Vo, Karina Kolmogortseva, Ngumimi Karen Iyortsuun, Hong-Duyen Vo, Soo-Hyung Kim

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11330 2025-10-14 cs.SD cs.AI cs.CL cs.LG eess.AS 67%

Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap

KiHyun Nam, Jongmin Choi, Hyeongkeun Lee, Jungwoo Heo, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology, South Korea(韩国科学技术院) University of Seoul, South Korea(首尔大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments 5 pages. Submitted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10201 2025-10-14 cs.LG cs.AI cs.CL 62%

RLFR: Extending Reinforcement Learning for LLMs with Flow Environment

Jinghao Zhang, Naishan Zheng, Ruilin Li, Dongzhou Cheng, Zheming Liang, Feng Zhao, Jiaqi Wang

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) ByteDance(字节跳动) Wuhan University(武汉大学) Southeast University(东南大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Project Website: https://jinghaoleven.github.io/RLFR/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11454 2025-10-14 cs.SD cs.AI 57%

Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning

Kuan-Yi Lee, Tsung-En Lin, Hung-Yi Lee

机构 * National Taiwan University(国立台湾大学) ASUS Open Cloud Infrastructure Software Center(ASUS开放云基础设施软件中心)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 9pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10329 2025-10-14 cs.CL 57%

End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs

Nam Luu, Ondřej Bojar

机构 * Charles University(查尔斯大学) Faculty of Mathematics and Physics(数学与物理系) Institute of Formal and Applied Linguistics(形式与应用语言学研究所)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10173 2025-10-14 cs.HC cs.CY cs.SD eess.AS 57%

Chord Colourizer: A Near Real-Time System for Visualizing Musical Key

Paul Haimes

机构 * Ritsumeikan University(立命馆大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Author copy. This paper is in press for presentation at ADADA 2025. Please cite as: Haimes, P. (in press). Chord Colourizer: A near real-time system for visualizing musical key. In Proceedings of the 23rd International Conference of Asia Digital Art and Design Association (ADADA)

详情

展开后加载摘要…

URL PDF HTML 收藏