arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-18 至 2025-11-18 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 9 篇

2511.12404 2025-11-18 cs.MM cs.AI cs.SD 81%

SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs

Shail Desai, Aditya Pawar, Li Lin, Xin Wang, Shu Hu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11930 2025-11-18 cs.HC cs.CV cs.LG cs.SD 79%

Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering

Tianyu Xu, Jihan Li, Penghe Zu, Pranav Sahay, Maruchi Kim, Jack Obeng-Marnu, Farley Miller, Xun Qian, Katrina Passarella, Mahitha Rachumalla, Rajeev Nongpiur, D. Shin

机构 * Google(谷歌)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Journal ref Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST '25), Article 17, 1-16, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10011 2025-11-18 cs.CY 78%

Reinforcing Trustworthiness in Multimodal Emotional Support Systems

Huy M. Le, Dat Tien Nguyen, Ngan T. T. Vo, Tuan D. Q. Nguyen, Nguyen Binh Le, Duy Minh Ho Nguyen, Daniel Sonntag, Lizi Liao, Binh T. Nguyen

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12452 2025-11-18 cs.CV cs.CL 73%

DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions

Xiaoyu Lin, Aniket Ghorpade, Hansheng Zhu, Justin Qiu, Dea Rrozhani, Monica Lama, Mick Yang, Zixuan Bian, Ruohan Ren, Alan B. Hong, Jiatao Gu, Chris Callison-Burch

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 音频语音多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11000 2025-11-18 cs.SD cs.AI 70%

DialogGraph-LLM: Graph-Informed LLMs for End-to-End Audio Dialogue Intent Recognition

HongYu Liu, Junxin Li, Changxi Guo, Hao Chen, Yaqian Huang, Yifu Guo, Huan Yang, Lihua Cai

机构 * South China Normal University, Guangzhou, China(华南师范大学) Xiamen Rekey Medical Technology Co., LTD, Xiamen, China(厦门瑞康医疗科技有限公司)

专题命中 音频语音多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

Comments 8 pages, 2 figures. To appear in: Proceedings of the 28th European Conference on Artificial Intelligence (ECAI 2025), Frontiers in Artificial Intelligence and Applications, Vol. 413. DOI: 10.3233/FAIA251182

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11634 2025-11-18 cs.RO cs.CV cs.HC cs.LG cs.MM 62%

Tactile Data Recording System for Clothing with Motion-Controlled Robotic Sliding

Michikuni Eguchi, Takekazu Kitagishi, Yuichi Hiroi, Takefumi Hiraki

机构 * University of Tsukuba(茨口大学) Cluster Metaverse Lab(元宇宙集群实验室) The University of Tokyo(东京大学) ZOZO Research(ZOZO研究所)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.MM

Comments 3 pages, 2 figures, 1 table. Presented at SIGGRAPH Asia 2025 Posters (SA Posters '25), December 15-18, 2025, Hong Kong, Hong Kong

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12662 2025-11-18 cs.CV 57%

Hi-Reco: High-Fidelity Real-Time Conversational Digital Humans

Hongbin Huang, Junwei Li, Tianxin Xie, Zhuang Li, Cekai Weng, Yaodong Yang, Yue Luo, Li Liu, Jing Tang, Zhijing Shao, Zeyu Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Prometheus Vision Technology Co., Ltd.(普罗米修斯视觉科技有限公司) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Proceedings of the Computer Graphics International 2025 (CGI'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08230 2025-11-18 cs.CL 57%

VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context

Heyang Liu, Ziyang Cheng, Yuhao Wang, Hongcheng Liu, Yiqi Li, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

Comments This article will serve as an extension of the preceding work, "VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models" (arXiv:2505.15727). Therefore, we have chosen to withdraw to avoid potential duplicate publication. We will update the previously open-sourced paper of VocalBench in several weeks to include the content of VocalBench-zh

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23447 2025-11-18 cs.CV 57%

CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition

Jongseo Lee, Joohyun Chang, Dongho Lee, Jinwoo Choi

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Our paper has been accepted to IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏