arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4587 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4587 篇

2502.21154 2025-08-21 cs.HC 78%

Hypergraph Multi-Modal Learning for EEG-based Emotion Recognition in Conversation

Zijian Kang, Yueyang Li, Shengyu Gong, Weiming Zeng, Hongjie Yan, Lingbin Bian, Zhiguo Zhang, Wai Ting Siok, Nizhuan Wang

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09402 2025-08-14 cs.HC 78%

Realtime Multimodal Emotion Estimation using Behavioral and Neurophysiological Data

Von Ralph Dane Marquez Herbuela, Yukie Nagai

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03300 2025-08-13 cs.RO cs.LG 78%

Touch and Tell: Multimodal Decoding of Human Emotions and Social Gestures for Robots

Qiaoqiao Ren, Remko Proesmans, Yuanbo Hou, Francis wyffels, Tony Belpaeme

机构 * Faculty of Engineering and Architecture(工程与建筑学院) IDLab-AIRO, Ghent University – imec(IDLab-AIRO,根特大学–imec) Department of Engineering Science, University of Oxford(工程科学系,牛津大学)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06117 2025-08-11 cs.HC 78%

A Multimodal Framework for Understanding Collaborative Design Processes

Maurice Koch, Nelusa Pathmanathan, Daniel Weiskopf, Kuno Kurzhals

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted to IEEE VIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04906 2025-07-31 cs.MM cs.CV cs.SD eess.AS 78%

Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation

Ivan Rinaldi, Nicola Fanelli, Giovanna Castellano, Gennaro Vessio

机构 * Department of Computer Science, University of Bari Aldo Moro, Italy(巴里阿尔多·莫罗大学计算机科学系)

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CV、cs.MM、eess.AS

Comments Presented at the AI for Visual Arts (AI4VA) workshop at ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15826 2025-07-22 cs.IR cs.LG 78%

Just Ask for Music (JAM): Multimodal and Personalized Natural Language Music Recommendation

Alessandro B. Melchiorre, Elena V. Epure, Shahed Masoudian, Gustavo Escobedo, Anna Hausberger, Manuel Moussallam, Markus Schedl

机构 * Johannes Kepler University Linz(约翰· Kepler大学林茨) Criteo AI Lab(Criteo人工智能实验室) Deezer Research(Deezer研究) Johannes Kepler University Linz and Linz Institute of Technology(约翰· Kepler大学林茨和林茨技术学院)

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01973 2025-07-15 cs.CE 78%

Multimodal Financial Foundation Models (MFFMs): Progress, Prospects, and Challenges

Xiao-Yang Liu Yanglet, Yupeng Cao, Li Deng

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14445 2025-06-18 cs.IR 78%

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

Ruofan Hu, Yan Xia, Minjie Hong, Jieming Zhu, Bo Chen, Xiaoda Yang, Minghui Fang, Tao Jin

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted by Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07598 2025-06-16 cs.HC 78%

Towards spatial computing: recent advances in multimodal natural interaction for XR headsets

Zhimin Wang, Maohang Rao, Shanghua Ye, Weitao Song, Feng Lu

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 28 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03378 2025-06-05 eess.AS cs.CV cs.MM 78%

SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer

Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera, Abu Osama Siddiqui, Sarthak Jain, Priyabrata Mallick, Jaya Sai Kiran Patibandla, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma

机构 * IIIT-DelhiIndia(印度德里印度理工学院) V.B.S.P.UIndia(印度V.B.S.P.U) UPESIndia(印度UPES) Independent ResearcherIndia(印度独立研究者) Reliance AIIndia(印度Reliance AI) University of TartuEstonia(爱沙尼亚塔尔图大学) Plaksha UniversityIndia(印度Plaksha大学)

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.MM、eess.AS

Comments Accepted to INTERSPEECH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13338 2025-06-04 cs.CL cs.AI eess.AS 78%

Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation

Qiongqiong Wang, Hardik B. Sailor, Tianchi Liu, Ai Ti Aw

机构 * Agency for Science, Technology and Research (A ⋆ ⋆ \star ⋆ STAR)(科技研究局) Institute for Infocomm Research (I 2 R)(信息通信研究所)

专题命中 音频语音多模态 :multi-modal(title);分类 cs.CL、cs.AI、eess.AS

Comments Accepted at Interspeech 2025. [v2]: The dataset has been released, and the link is now updated

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19877 2025-05-05 cs.HC 78%

Towards Multimodal Large-Language Models for Parent-Child Interaction: A Focus on Joint Attention

Weiyan Shi, Viet Hai Le, Kenny Tsu Wei Choo

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted at CHI 2025 Late Breaking Work

Journal ref Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 2025, Article No. 535, Pages 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18914 2025-04-29 cs.LG stat.AP stat.ML 78%

Factor Analysis with Correlated Topic Model for Multi-Modal Data

Małgorzata Łazęcka, Ewa Szczurek

机构 * Faculty of Mathematics, Informatics and Mechanics, University of Warsaw(华沙大学数学、信息学与力学学院) Institute of Computer Science, Polish Academy of Sciences(波兰科学院计算机科学研究所) Institute of AI for Health, Helmholtz Center Munich(慕尼黑海德堡中心人工智能与健康研究所)

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract)

Comments AISTATS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14214 2025-04-22 cs.IR 78%

Teach Me How to Denoise: A Universal Framework for Denoising Multi-modal Recommender Systems via Guided Calibration

Hongji Li, Hanwen Du, Youhua Li, Junchen Fu, Chunxiao Li, Ziyi Zhuang, Jiakang Li, Yongxin Ni

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments Accepted to ACM Web Search and Data Mining (WSDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11460 2025-04-21 cs.CV cs.AI cs.CL 78%

Semantic Matters: Multimodal Features for Affective Analysis

Tobias Hallmen, Robin-Nico Kampa, Fabian Deuser, Norbert Oswald, Elisabeth André

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06593 2025-04-10 cs.RO cs.HC 78%

A Multi-Modal Interaction Framework for Efficient Human-Robot Collaborative Shelf Picking

Abhinav Pathak, Kalaichelvi Venkatesan, Tarek Taha, Rajkumar Muthusamy

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00785 2025-04-07 cs.RO 78%

Natural Multimodal Fusion-Based Human-Robot Interaction: Application With Voice and Deictic Posture via Large Language Model

Yuzhi Lai, Shenghai Yuan, Youssef Nassar, Mingyu Fan, Atmaraaj Gopal, Arihiro Yorita, Naoyuki Kubota, Matthias Rätsch

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract)

Comments Accepted for publication by IEEE Robotics & Automation Magazine

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19334 2025-03-26 cs.HC 78%

Design of Seamless Multi-modal Interaction Framework for Intelligent Virtual Agents in Wearable Mixed Reality Environment

Ghazanfar Ali, Hong-Quan Le, Junho Kim, Seoung-won Hwang, Jae-In Hwang

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract)

Comments 6 pages, 14 Figures, Computer Animation and Social Agents (CASA 2019)

Journal ref CASA 2019: Proceedings of the 32nd International Conference on Computer Animation and Social Agents - Year 2019 - Pages 47 - 52

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01226 2025-03-04 q-bio.NC cs.LG 78%

Dementia Insights: A Context-Based MultiModal Approach

Sahar Sinene Mehdoui, Abdelhamid Bouzid, Daniel Sierra-Sosa, Adel Elmaghraby

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04592 2025-03-04 cs.HC 78%

CardioAI: A Multimodal AI-based System to Support Symptom Monitoring and Risk Detection of Cancer Treatment-Induced Cardiotoxicity

Siyi Wu, Weidan Cao, Shihan Fu, Bingsheng Yao, Ziqi Yang, Changchang Yin, Varun Mishra, Daniel Addison, Ping Zhang, Dakuo Wang

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17971 2025-02-26 cs.RO cs.HC 78%

Multimodal Interaction and Intention Communication for Industrial Robots

Tim Schreiter, Andrey Rudenko, Jens V. Rüppel, Martin Magnusson, Achim J. Lilienthal

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted to the 1st German Robotics Conference (GRC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16703 2025-02-18 cs.RO cs.SY eess.SP eess.SY 78%

RoboMNIST: A Multimodal Dataset for Multi-Robot Activity Recognition Using WiFi Sensing, Video, and Audio

Kian Behzad, Rojin Zandi, Elaheh Motamedi, Hojjat Salehinejad, Milad Siami

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18940 2025-02-17 cs.HC 78%

Amuse: Human-AI Collaborative Songwriting with Multimodal Inspirations

Yewon Kim, Sung-Ju Lee, Chris Donahue

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Published as a conference paper at CHI 2025. Project page: https://yewon-kim.com/amuse

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03862 2025-02-10 cs.HC 78%

Enhancing Deliberativeness: Evaluating the Impact of Multimodal Reflection Nudges

ShunYi Yeo, Zhuoqun Jiang, Anthony Tang, Simon Tangi Perrault

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02830 2025-02-06 cs.HC cs.LG q-bio.NC 78%

Multimodal Brain-Computer Interfaces: AI-powered Decoding Methodologies

Siyang Li, Hongbin Wang, Xiaoqing Chen, Dongrui Wu

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01801 2025-02-05 cs.HC 78%

MemPal: Leveraging Multimodal AI and LLMs for Voice-Activated Object Retrieval in Homes of Older Adults

Natasha Maniar, Samantha W. T. Chan, Wazeer Zulfikar, Scott Ren, Christine Xu, Pattie Maes

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15711 2025-02-05 cs.RO 78%

Robustifying Long-term Human-Robot Collaboration through a Multimodal and Hierarchical Framework

Peiqi Yu, Abulikemu Abuduweili, Ruixuan Liu, Changliu Liu

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12088 2025-02-05 cs.CY 78%

Mental-Perceiver: Audio-Textual Multi-Modal Learning for Estimating Mental Disorders

Jinghui Qin, Changsong Liu, Tianchi Tang, Dahuang Liu, Minghao Wang, Qianying Huang, Rumin Zhang

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments Accepted to AAAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05322 2025-01-14 cs.HC 78%

"What's Happening"- A Human-centered Multimodal Interpreter Explaining the Actions of Autonomous Vehicles

Xuewen Luo, Fan Ding, Ruiqi Chen, Rishikesh Panda, Junnyong Loo, Shuyun Zhang

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This paper has been accepted for presentation at WACV Workshop HAVI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12273 2024-12-31 cs.RO 78%

Multimodal Human-Autonomous Agents Interaction Using Pre-Trained Language and Visual Foundation Models

Linus Nwankwo, Elmar Rueckert

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏