arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2408.03813 2025-09-18 cs.HC 50%

Talk to the Wall: The Role of Speech Interaction in Collaborative Visual Analytics

Gabriela Molina León, Anastasia Bezerianos, Olivier Gladin, Petra Isenberg

专题命中 音频语音多模态 :multimodal(abstract)

Comments 11 pages, 6 figures, to appear in IEEE TVCG (VIS 2024); correct figure

Journal ref IEEE Transactions on Visualization and Computer Graphics, 31(1), 2025, 941-951

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.02051 2025-09-16 cs.CR 50%

Sanitization of Multimedia Content: A Survey of Techniques, Attacks, and Future Directions

Andrea Ciccotelli, Hanaa Abbas, Roberto Di Pietro

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04392 2025-09-05 cs.SD 50%

Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition

Yanyan Liu, Minqiang Xu, Yihao Chen, Liang He, Lei Fang, Sian Fang, Lin Liu

机构 * School of Computer Science and Technology(计算机科学与技术学院) Xinjiang University(新疆大学) Hefei iFly Digital Technology Co. Ltd.(合肥iFly数字技术有限公司) University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学)

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04356 2025-09-05 cs.HC cs.RO 50%

SRWToolkit: An Open Source Wizard of Oz Toolkit to Create Social Robotic Avatars

Atikkhan Faridkhan Nilgar, Kristof Van Laerhoven, Ayub Kinoti

机构 * University of Siegen(施皮格恩大学) Honda Research Institute Europe GmbH(本田欧洲研究院) Dedan Kimathi University of Technology(德丹·基马蒂技术大学)

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref 2025 International Conference on Social Robotics (ICSR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01246 2025-09-03 cs.HC cs.RO 50%

An AI-Based Shopping Assistant System to Support the Visually Impaired

Larissa R. de S. Shibata, Ankit A. Ravankar, Jose Victorio Salazar Luces, Yasuhisa Hirata

机构 * Department of Robotics, Tohoku University(机器人系,东京东京大学)

专题命中 音频语音多模态 :multimodal(abstract)

Comments 7 pages, Accepted for 2025 SICE-FES conference (IEEE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19514 2025-08-28 cs.SD 50%

MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models

Zhihao Ouyang, Ju-Chiang Wang, Daiyu Zhang, Bin Chen, Shangjie Li, Quan Lin

机构 * ByteDance(字节跳动)

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17124 2025-08-26 cs.HC cs.SY eess.SY 50%

Towards Deeper Understanding of Natural User Interactions in Virtual Reality Based Assembly Tasks

Ryan Ghamandi, Yahya Hmaiti, Mykola Maslych, Ravi Kiran Kattoju, Joseph J. LaViola

专题命中 音频语音多模态 :multimodal(abstract)

Comments To be submitted in a future conference, this is the author version pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12918 2025-08-22 cs.SD 50%

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation

Lei Zhao, Rujin Chen, Chi Zhang, Xiao-Lei Zhang, Xuelong Li

机构 * School of Marine Science and Technology, Northwestern Polytechnical University(海洋科学与技术学院,西北工业大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) Research and Development Institute of Northwestern Polytechnical University in Shenzhen, China(西北工业大学深圳研发院,中国)

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12186 2025-08-19 cs.SI 50%

MAD: A Benchmark for Multi-Turn Audio Dialogue Fact-Checking

Chaewan Chun, Lysandre Terrisse, Delvin Ce Zhang, Dongwon Lee

专题命中 音频语音多模态 :multimodal(abstract)

Comments 11 pages, Accepted to SBP-BRiMS 2025 Working Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10130 2025-08-15 q-bio.NC 50%

Linking GFAP Levels to Speech Anomalies in Acute Brain Injury: A Simulation Based Study

Shamaley Aravinthan, Bin Hu

专题命中 音频语音多模态 :multimodal(abstract)

Comments 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01092 2025-08-05 cs.HC 50%

DescribePro: Collaborative Audio Description with Human-AI Interaction

Maryam Cheema, Sina Elahimanesh, Samuel Martin, Pooyan Fazli, Hasti Seifi

专题命中 音频语音多模态 :multimodal(abstract)

Comments ASSETS 25 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18478 2025-07-25 cs.CR 50%

Scout: Leveraging Large Language Models for Rapid Digital Evidence Discovery

Shariq Murtuza

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17405 2025-07-24 q-bio.NC 50%

Automatic Blink-based Bad EEG channels Detection for BCI Applications

Eva Guttmann-Flury, Yanyan Wei, Shan Zhao

专题命中 音频语音多模态 :multimodal(abstract)

Comments 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (IEEE EMBC 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16074 2025-07-23 cs.HC 50%

Toward music-based stress management: Contemporary biosensing systems for affective regulation

Natasha Yamane, Varun Mishra, Matthew S. Goodwin

专题命中 音频语音多模态 :multimodal(abstract)

Comments 37 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15729 2025-07-22 cs.RO cs.HC 50%

Gaze-supported Large Language Model Framework for Bi-directional Human-Robot Interaction

Jens V. Rüppel, Andrey Rudenko, Tim Schreiter, Martin Magnusson, Achim J. Lilienthal

机构 * Technical University of Munich(慕尼黑技术大学) Robert Bosch GmbH, Corporate Research(博世集团,企业研究) Centre for Applied Autonomous Sensor Systems (AASS)(应用自主传感器系统中心(AASS))

专题命中 音频语音多模态 :multi-modal(abstract)

Comments This paper has been accepted to the 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), which will be held in Eindhoven, Netherlands on August 25-29, 2025. Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06802 2025-07-10 cs.LG 50%

Speech Tokenizer is Key to Consistent Representation

Wonjin Jung, Sungil Kang, Dong-Yeon Cho

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13901 2025-07-03 cs.HC 50%

Examining Technology Perspectives of Older Adults with Mild Cognitive Impairment: A Scoping Review

Snezna B Schmidt, Stephen Isbel, Blooma John, Ram Subramanian, Nathan M DCunha

专题命中 音频语音多模态 :multimodal(abstract)

Comments Paper updated

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00672 2025-07-02 cs.NI cs.DC 50%

Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration

Haoxiang Luo, Yinqiu Liu, Ruichen Zhang, Jiacheng Wang, Gang Sun, Dusit Niyato, Hongfang Yu, Zehui Xiong, Xianbin Wang, Xuemin Shen

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23443 2025-07-01 cs.HC 50%

Accessible Data Access and Analysis by People who are Blind or Have Low Vision

Samuel Reinders, Munazza Zaib, Matthew Butler, Bongshin Lee, Ingrid Zukerman, Lizhen Qu, Kim Marriott

专题命中 音频语音多模态 :multimodal(abstract)

Comments Poster presented at the 1st Workshop on Accessible Data Visualization, IEEE VIS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21983 2025-06-30 eess.SP 50%

Learning-Based Hybrid Neural Receiver for 6G-V2X Communications

Osama Saleem, Mohammed Alfaqawi, Pierre Merdrignac, Abdelaziz Bensrhair, Soheyb Ribouh

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16038 2025-06-26 cs.SI 50%

Emotion-Aware Design: Modulating Valence, Arousal, and Dominance in Communication via Design

Shixiong Cao, Nan Cao

专题命中 音频语音多模态 :multimodal(abstract)

Comments 17 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18711 2025-06-24 cs.HC 50%

LLM-enhanced Interactions in Human-Robot Collaborative Drawing with Older Adults

Marianne Bossema, Somaya Ben Allouch, Aske Plaat, Rob Saunders

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17890 2025-06-24 cs.HC 50%

One Does Not Simply 'Mm-hmm': Exploring Backchanneling in the AAC Micro-Culture

Tobias Weinberg, Claire O'Connor, Ricardo E. Gonzalez Penuela, Stephanie Valencia, Thijs Roumen

专题命中 音频语音多模态 :multi-modal(abstract)

Comments See our project and video at: https://tobiwg.com/research/one_does_not_simply_hm-hmm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16716 2025-06-23 cs.HC 50%

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos

Qixin Wang, Songtao Zhou, Zeyu Jin, Chenglin Guo, Shikun Sun, Xiaoyu Qin

专题命中 音频语音多模态 :audio-visual(abstract)

Comments Accepted by IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15497 2025-06-19 cs.HC 50%

Foundation of Affective Computing and Interaction

Changzeng Fu

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11808 2025-06-18 cs.LG stat.ME stat.ML 50%

Understanding the Trade-offs in Accuracy and Uncertainty Quantification: Architecture and Inference Choices in Bayesian Neural Networks

Alisa Sheinkman, Sara Wade

机构 * School of Mathematics and Maxwell Institute for Mathematical Sciences, University of Edinburgh(数学系和Maxwell数学科学研究所,爱丁堡大学)

专题命中 音频语音多模态 :multimodal(abstract)

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10930 2025-06-13 cs.LG 50%

Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction

Thanathai Lertpetchpun, Tiantian Feng, Dani Byrd, Shrikanth Narayanan

机构 * University of Southern California(南加州大学)

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref Accepeted to Interspeech2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09542 2025-06-13 q-bio.NC stat.ME 50%

Multi-object Data Integration in the Study of Primary Progressive Aphasia

Rene Gutierrez, Rajarshi Guhaniyogi, Aaron Scheffler, Maria Luisa Gorno-Tempini, Maria Luisa Mandelli, Giovanni Battistella

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23732 2025-05-30 cs.LG 50%

EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast

Shreeram Suresh Chandra, Lucas Goncalves, Junchen Lu, Carlos Busso, Berrak Sisman

机构 * Center for Language and Speech Processing (CLSP), Johns Hopkins UniversityUSA(语言与语音处理中心(CLSP),约翰霍普金斯大学) The University of Texas at Dallas, USA(德克萨斯大学达拉斯分校) Amazon, USA(亚马逊) NUSSingapore(南洋理工大学) Language Technologies Institute (LTI), Carnegie Mellon UniversityUSA(语言技术研究所(LTI),卡内基梅隆大学)

专题命中 音频语音多模态 :cross-modal(abstract)

Comments Accepted at Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11159 2025-05-19 quant-ph physics.ed-ph 50%

Sonification of entanglement dynamics in many-qubit systems

Juliette Tudoce, Marcin Płodzień, Maciej Lewenstein, Reiko Yamada

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏