arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-25 至 2025-08-25 共收录 7 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 7 篇

2508.16143 2025-08-25 cs.RO cs.AI 79%

Take That for Me: Multimodal Exophora Resolution with Interactive Questioning for Ambiguous Out-of-View Instructions

Akira Oyama, Shoichi Hasegawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi

机构 * Ritsumeikan University(立命馆大学) Soka University(早稻田大学) Kyoto University(京都大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments See website at https://emergentsystemlabstudent.github.io/MIEL/. Accepted at IEEE RO-MAN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16320 2025-08-25 physics.ed-ph 78%

AI-Supported Mini-Labs: Combining Smartphone-Based Experiments and Multimodal AI

Jochen Kuhn, David J. Rakestraw, Stefan Küchemann, Patrik Vogt

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15810 2025-08-25 cs.CL cs.AI cs.LG 76%

Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models

Nouar AlDahoul, Yasir Zaki

机构 * Computer Science Department, New York University Abu Dhabi(纽约大学阿布扎克分校计算机科学系)

专题命中 音频语音多模态 :multi-modal(title);分类 cs.CL、cs.AI

Comments 26 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15716 2025-08-25 cs.HC cs.AI 70%

Foundation Models for Cross-Domain EEG Analysis Application: A Survey

Hongqi Li, Yitong Chen, Yujuan Wang, Weihang Ni, Haodong Zhang

机构 * School of Software, Northwestern Polytechnical University(软件学院,西北工业大学)

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments Submitted to IEEE Journals

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.21054 2025-08-25 cs.CL cs.AI cs.LG cs.SD eess.AS 67%

Sentiment Reasoning for Healthcare

Khai-Nguyen Nguyen, Khai Le-Duc, Bach Phan Tat, Duy Le, Long Vo-Dang, Truong-Son Hy

机构 * College of William and Mary(威廉与玛丽学院) University of Toronto(多伦多大学) University Health Network(大学健康网络) KU Leuven(鲁汶大学) Bucknell University(巴克内尔大学) University of Cincinnati(辛辛那提大学) University of Alabama at Birmingham(伯明翰大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments ACL 2025 Industry Track (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16835 2025-08-25 eess.AS cs.CL 62%

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems

Rumi Allbert, Nima Yazdani, Ali Ansari, Aruj Mahajan, Amirhossein Afsharrad, Seyed Shahabeddin Mousavi

机构 * University of Southern California(南加州大学) Stanford University(斯坦福大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00852 2025-08-25 cs.HC cs.CV cs.LG cs.RO 57%

Visuo-Acoustic Hand Pose and Contact Estimation

Yuemin Mao, Uksang Yoo, Yunchao Yao, Shahram Najam Syed, Luca Bondi, Jonathan Francis, Jean Oh, Jeffrey Ichnowski

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏