arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-01 至 2025-10-01 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 9 篇

2505.20873 2025-10-01 cs.CV 88%

Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models

Chaeyoung Jung, Youngjoon Jang, Jongmin Choi, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20862 2025-10-01 cs.CV 85%

AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding

Chaeyoung Jung, Youngjoon Jang, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25652 2025-10-01 cs.AI cs.MM cs.SD 84%

Iterative Residual Cross-Attention Mechanism: An Integrated Approach for Audio-Visual Navigation Tasks

Hailong Zhang, Yinfeng Yu, Liejun Wang, Fuchun Sun, Wendong Zheng

机构 * Xinjiang Multimodal Intelligent Processing and Information Security Engineering Technology Research Center, School of Computer Science and Technology, Xinjiang University(新疆多模态智能处理与信息安全工程技术创新中心,计算机科学与技术学院,新疆大学) Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学) School of Electrical Engineering and Automation, Tianjin University of Technology(电气工程与自动化学院,天津工业大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.AI、cs.MM

Comments Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15772 2025-10-01 cs.SD cs.CL eess.AS 84%

MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling

Yifan Cheng, Ruoyi Zhang, Jiatong Shi

机构 * Santa Clara, CA, USA(美国圣克拉拉) Carnegie Mellon University(卡内基梅隆大学) Huazhong University of Science and Technology(华中科技大学) Nanjing University of Information Science and Technology(南京信息工程大学)

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CL、eess.AS

Comments Accepted by Interspeech

Journal ref Proc. of Interspeech2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13624 2025-10-01 cs.SD eess.AS 79%

Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement

Rong Chao, Wenze Ren, You-Jin Li, Kuo-Hsuan Hung, Sung-Feng Huang, Szu-Wei Fu, Wen-Huang Cheng, Yu Tsao

机构 * Academia Sinica(台湾“中央研究院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted to Interspeech 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26604 2025-10-01 cs.CV 57%

Video Object Segmentation-Aware Audio Generation

Ilpo Viertola, Vladimir Iashin, Esa Rahtu

机构 * Tampere University, Tampere, Finland(塔尔库大学) University of Oxford, Oxford, UK(牛津大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Preprint version. The Version of Record is published in DAGM GCPR 2025 proceedings with Springer Lecture Notes in Computer Science (LNCS). Updated results and resources are available at the project page: https://saganet.notion.site

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25670 2025-10-01 cs.SD cs.CV 57%

LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning

Kang Yang, Yifan Liang, Fangkun Liu, Zhenping Xie, Chengshi Zheng

机构 * School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) Institute of Acoustics Chinese Academy of Science(中国科学院声学研究所)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26593 2025-10-01 cs.HC 50%

Exploring Large Language Model as an Interactive Sports Coach: Lessons from a Single-Subject Half Marathon Preparation

Kichang Lee

专题命中 音频语音多模态 :multimodal(abstract)

Comments 23 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14608 2025-10-01 cs.RO 50%

Visual-auditory Extrinsic Contact Estimation

Xili Yi, Jayjun Lee, Nima Fazeli

机构 * Robotics Department, University of Michigan(密歇根大学机器人系)

专题命中 音频语音多模态 :multimodal(abstract)

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏