arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2508.12918 2025-08-22 cs.SD 50%

FoleySpace: Vision-Aligned Binaural Spatial Audio Generation

Lei Zhao, Rujin Chen, Chi Zhang, Xiao-Lei Zhang, Xuelong Li

机构 * School of Marine Science and Technology, Northwestern Polytechnical University(海洋科学与技术学院,西北工业大学) Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) Research and Development Institute of Northwestern Polytechnical University in Shenzhen, China(西北工业大学深圳研发院,中国)

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12186 2025-08-19 cs.SI 50%

MAD: A Benchmark for Multi-Turn Audio Dialogue Fact-Checking

Chaewan Chun, Lysandre Terrisse, Delvin Ce Zhang, Dongwon Lee

专题命中 音频语音多模态 :multimodal(abstract)

Comments 11 pages, Accepted to SBP-BRiMS 2025 Working Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10130 2025-08-15 q-bio.NC 50%

Linking GFAP Levels to Speech Anomalies in Acute Brain Injury: A Simulation Based Study

Shamaley Aravinthan, Bin Hu

专题命中 音频语音多模态 :multimodal(abstract)

Comments 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01092 2025-08-05 cs.HC 50%

DescribePro: Collaborative Audio Description with Human-AI Interaction

Maryam Cheema, Sina Elahimanesh, Samuel Martin, Pooyan Fazli, Hasti Seifi

专题命中 音频语音多模态 :multimodal(abstract)

Comments ASSETS 25 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18478 2025-07-25 cs.CR 50%

Scout: Leveraging Large Language Models for Rapid Digital Evidence Discovery

Shariq Murtuza

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17405 2025-07-24 q-bio.NC 50%

Automatic Blink-based Bad EEG channels Detection for BCI Applications

Eva Guttmann-Flury, Yanyan Wei, Shan Zhao

专题命中 音频语音多模态 :multimodal(abstract)

Comments 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (IEEE EMBC 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16074 2025-07-23 cs.HC 50%

Toward music-based stress management: Contemporary biosensing systems for affective regulation

Natasha Yamane, Varun Mishra, Matthew S. Goodwin

专题命中 音频语音多模态 :multimodal(abstract)

Comments 37 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15729 2025-07-22 cs.RO cs.HC 50%

Gaze-supported Large Language Model Framework for Bi-directional Human-Robot Interaction

Jens V. Rüppel, Andrey Rudenko, Tim Schreiter, Martin Magnusson, Achim J. Lilienthal

机构 * Technical University of Munich(慕尼黑技术大学) Robert Bosch GmbH, Corporate Research(博世集团,企业研究) Centre for Applied Autonomous Sensor Systems (AASS)(应用自主传感器系统中心(AASS))

专题命中 音频语音多模态 :multi-modal(abstract)

Comments This paper has been accepted to the 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), which will be held in Eindhoven, Netherlands on August 25-29, 2025. Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06802 2025-07-10 cs.LG 50%

Speech Tokenizer is Key to Consistent Representation

Wonjin Jung, Sungil Kang, Dong-Yeon Cho

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13901 2025-07-03 cs.HC 50%

Examining Technology Perspectives of Older Adults with Mild Cognitive Impairment: A Scoping Review

Snezna B Schmidt, Stephen Isbel, Blooma John, Ram Subramanian, Nathan M DCunha

专题命中 音频语音多模态 :multimodal(abstract)

Comments Paper updated

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00672 2025-07-02 cs.NI cs.DC 50%

Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration

Haoxiang Luo, Yinqiu Liu, Ruichen Zhang, Jiacheng Wang, Gang Sun, Dusit Niyato, Hongfang Yu, Zehui Xiong, Xianbin Wang, Xuemin Shen

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23443 2025-07-01 cs.HC 50%

Accessible Data Access and Analysis by People who are Blind or Have Low Vision

Samuel Reinders, Munazza Zaib, Matthew Butler, Bongshin Lee, Ingrid Zukerman, Lizhen Qu, Kim Marriott

专题命中 音频语音多模态 :multimodal(abstract)

Comments Poster presented at the 1st Workshop on Accessible Data Visualization, IEEE VIS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21983 2025-06-30 eess.SP 50%

Learning-Based Hybrid Neural Receiver for 6G-V2X Communications

Osama Saleem, Mohammed Alfaqawi, Pierre Merdrignac, Abdelaziz Bensrhair, Soheyb Ribouh

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16038 2025-06-26 cs.SI 50%

Emotion-Aware Design: Modulating Valence, Arousal, and Dominance in Communication via Design

Shixiong Cao, Nan Cao

专题命中 音频语音多模态 :multimodal(abstract)

Comments 17 pages,6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18711 2025-06-24 cs.HC 50%

LLM-enhanced Interactions in Human-Robot Collaborative Drawing with Older Adults

Marianne Bossema, Somaya Ben Allouch, Aske Plaat, Rob Saunders

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17890 2025-06-24 cs.HC 50%

One Does Not Simply 'Mm-hmm': Exploring Backchanneling in the AAC Micro-Culture

Tobias Weinberg, Claire O'Connor, Ricardo E. Gonzalez Penuela, Stephanie Valencia, Thijs Roumen

专题命中 音频语音多模态 :multi-modal(abstract)

Comments See our project and video at: https://tobiwg.com/research/one_does_not_simply_hm-hmm/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16716 2025-06-23 cs.HC 50%

V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos

Qixin Wang, Songtao Zhou, Zeyu Jin, Chenglin Guo, Shikun Sun, Xiaoyu Qin

专题命中 音频语音多模态 :audio-visual(abstract)

Comments Accepted by IJCNN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15497 2025-06-19 cs.HC 50%

Foundation of Affective Computing and Interaction

Changzeng Fu

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11808 2025-06-18 cs.LG stat.ME stat.ML 50%

Understanding the Trade-offs in Accuracy and Uncertainty Quantification: Architecture and Inference Choices in Bayesian Neural Networks

Alisa Sheinkman, Sara Wade

机构 * School of Mathematics and Maxwell Institute for Mathematical Sciences, University of Edinburgh(数学系和Maxwell数学科学研究所,爱丁堡大学)

专题命中 音频语音多模态 :multimodal(abstract)

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10930 2025-06-13 cs.LG 50%

Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction

Thanathai Lertpetchpun, Tiantian Feng, Dani Byrd, Shrikanth Narayanan

机构 * University of Southern California(南加州大学)

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref Accepeted to Interspeech2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09542 2025-06-13 q-bio.NC stat.ME 50%

Multi-object Data Integration in the Study of Primary Progressive Aphasia

Rene Gutierrez, Rajarshi Guhaniyogi, Aaron Scheffler, Maria Luisa Gorno-Tempini, Maria Luisa Mandelli, Giovanni Battistella

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23732 2025-05-30 cs.LG 50%

EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast

Shreeram Suresh Chandra, Lucas Goncalves, Junchen Lu, Carlos Busso, Berrak Sisman

机构 * Center for Language and Speech Processing (CLSP), Johns Hopkins UniversityUSA(语言与语音处理中心(CLSP),约翰霍普金斯大学) The University of Texas at Dallas, USA(德克萨斯大学达拉斯分校) Amazon, USA(亚马逊) NUSSingapore(南洋理工大学) Language Technologies Institute (LTI), Carnegie Mellon UniversityUSA(语言技术研究所(LTI),卡内基梅隆大学)

专题命中 音频语音多模态 :cross-modal(abstract)

Comments Accepted at Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11159 2025-05-19 quant-ph physics.ed-ph 50%

Sonification of entanglement dynamics in many-qubit systems

Juliette Tudoce, Marcin Płodzień, Maciej Lewenstein, Reiko Yamada

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05441 2025-05-09 cs.HC 50%

GesPrompt: Leveraging Co-Speech Gestures to Augment LLM-Based Interaction in Virtual Reality

Xiyun Hu, Dizhi Ma, Fengming He, Zhengzhe Zhu, Shao-Kang Hsia, Chenfei Zhu, Ziyi Liu, Karthik Ramani

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00153 2025-05-02 cs.HC cs.DC 50%

Audo-Sight: Enabling Ambient Interaction For Blind And Visually Impaired Individuals

Bhanuja Ainary

专题命中 音频语音多模态 :multimodal(abstract)

Comments This thesis was conducted under the guidance of Mohsen Amini Salehi. Special thanks to Minseo Kim and Jacob Bradshaw for their valuable contributions and support throughout the research process. 60 pages, 13 Figures, 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09435 2025-04-15 cs.HC 50%

Design Probes for AI-Driven AAC: Addressing Complex Communication Needs in Aphasia

Lei Mao, Jong Ho Lee, Yasmeen Faroqi Shah, Stephanie Valencia

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06379 2025-04-10 cs.HC cs.SY eess.SY 50%

User-Centered Insights into Assistive Navigation Technologies for Individuals with Visual Impairment

Iman Soltani, Johnaton Schofield, Mehran Madani, Daniel Kish, Parisa Emami-Naeini

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04388 2025-04-08 cs.CR 50%

Who's Watching You Zoom? Investigating Privacy of Third-Party Zoom Apps

Saharsh Goenka, Adit Prabhu, Payge Sakurai, Mrinaal Ramachandran, Rakibul Hasan

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01708 2025-04-03 cs.RO cs.HC cs.LG 50%

TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication

Petr Vanc, Karla Stepanova

机构 * Czech Technical University in Prague(布拉格捷克技术大学) Czech Institute of Informatics, Robotics, and Cybernetics(捷克信息学、机器人学与控制论研究所)

专题命中 音频语音多模态 :multimodal(abstract)

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19398 2025-03-26 cs.HC 50%

CyanKitten: AI-Driven Markerless Motion Capture for Improved Elderly Well-Being

Mengyao Guo, Yu Nie, Jinda Han, Zongxing Li, Ze Gao

专题命中 音频语音多模态 :multimodal(abstract)

Comments Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, April 26-May 1, 2025, Yokohama, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏