arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-18 至 2025-09-18 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 7 篇

2509.14097 2025-09-18 cs.CV cs.MM 88%

Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing

Yaru Chen, Ruohao Guo, Liting Gao, Yang Xiang, Qingyu Luo, Zhenbo Li, Wenwu Wang

专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(title,abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09595 2025-09-18 cs.CV 85%

Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis

Yikang Ding, Jiwen Liu, Wenyuan Zhang, Zekun Wang, Wentao Hu, Liyuan Cui, Mingming Lao, Yingchao Shao, Hui Liu, Xiaohan Li, Ming Chen, Xiaoqiang Liu, Yu-Shen Liu, Pengfei Wan

机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);audio-visual(abstract);分类 cs.CV

Comments Technical Report. Project Page: https://klingavatar.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13395 2025-09-18 eess.AS cs.AI cs.CL cs.LG cs.MM 83%

TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models

Haolong Zheng, Yekaterina Yegorova, Mark Hasegawa-Johnson

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15871 2025-09-18 cs.CY cs.AI cs.CL 62%

A Comprehensive Survey on the Trustworthiness of Large Language Models in Healthcare

Manar Aljohani, Jun Hou, Sindhura Kommu, Xuan Wang

机构 * Virginia Tech(维吉尼亚理工大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14023 2025-09-18 cs.CL cs.HC 57%

Audio-Based Crowd-Sourced Evaluation of Machine Translation Quality

Sami Ul Haq, Sheila Castilho, Yvette Graham

机构 * ADAPT Centre(ADAPT中心) Dublin City University (DCU)(都柏林城市大学) Trinity College Dublin (TCD)(三一学院都柏林)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted at WMT2025 (ENNLP) for oral presented

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10432 2025-09-18 q-bio.OT cs.AI 57%

Standards in the Preparation of Biomedical Research Metadata: A Bridge2AI Perspective

Harry Caufield, Satrajit Ghosh, Sek Wong Kong, Jillian Parker, Nathan Sheffield, Bhavesh Patel, Andrew Williams, Timothy Clark, Monica C. Munoz-Torres

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03813 2025-09-18 cs.HC 50%

Talk to the Wall: The Role of Speech Interaction in Collaborative Visual Analytics

Gabriela Molina León, Anastasia Bezerianos, Olivier Gladin, Petra Isenberg

专题命中 音频语音多模态 :multimodal(abstract)

Comments 11 pages, 6 figures, to appear in IEEE TVCG (VIS 2024); correct figure

Journal ref IEEE Transactions on Visualization and Computer Graphics, 31(1), 2025, 941-951

详情

展开后加载摘要…

URL PDF HTML 收藏