arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4587 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4587 篇

2412.00185 2024-12-03 astro-ph.IM astro-ph.GA 78%

Interactive Multimodal Integral Field Spectroscopy

Adrián García Riber, Rubén García-Benito, Francisco Serradilla

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 12 pages, 12 figures, 2 tables. Accepted for publication in RAS Techniques & Instruments (RASTI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18587 2024-11-28 cs.HC eess.SP q-bio.NC 78%

EEG-Based Analysis of Brain Responses in Multi-Modal Human-Robot Interaction: Modulating Engagement

Suzanne Oliver, Tomoko Kitago, Adam Buchwald, S. Farokh Atashzar

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments 9 pages, 7 figures. Submitted to IEEE TNSRE

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15590 2024-11-26 cs.LG cs.HC stat.ME 78%

From Complexity to Parsimony: Integrating Latent Class Analysis to Uncover Multimodal Learning Patterns in Collaborative Learning

Lixiang Yan, Dragan Gašević, Linxuan Zhao, Vanessa Echeverria, Yueqiao Jin, Roberto Martinez-Maldonado

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.14147 2024-11-22 eess.IV 78%

Spiking neural networks: Towards bio-inspired multimodal perception in robotics

Katerina Maria Oikonomou, Vasiliki Balaska, Konstantinos A. Tsintotas, Christos N. Mavridis, Ioannis Kansizoglou, Antonios Gasteratos

专题命中 音频语音多模态 :multimodal(title);audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05342 2024-11-11 cs.RO 78%

Development of a Human-Robot Interaction Platform for Dual-Arm Robots Based on ROS and Multimodal Artificial Intelligence

Thanh Nguyen Canh, Ba Phuong Nguyen, Hong Quan Tran, Xiem HoangVan

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments In The 25th National Conference on Electronics, Communications and Information Technology (REV-ECIT 2022), Hanoi, Vietnam. in Vietnamese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19464 2024-11-05 cs.RO cs.AI cs.CV cs.SD eess.AS 78%

ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data

Zeyi Liu, Cheng Chi, Eric Cousineau, Naveen Kuppuswamy, Benjamin Burchfiel, Shuran Song

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV、cs.AI、eess.AS

Comments Conference on Robot Learning (CoRL) 2024; Project website: https://maniwav.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04279 2024-10-28 cs.SI 78%

MSEVA : A System for Multimodal Short Videos Emotion Visual Analysis

Qinglan Wei, Yaqi Zhou, Longhui Xiao, Yuan Zhang

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.01072 2024-10-28 cs.LG 78%

Interpretability for Multimodal Emotion Recognition using Concept Activation Vectors

Ashish Ramayee Asokan, Nidarshan Kumar, Anirudh Venkata Ragam, Shylaja S Sharath

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08578 2024-10-22 cs.HC eess.SP 78%

Dynamics of Collective Group Affect: Group-level Annotations and the Multimodal Modeling of Convergence and Divergence

Navin Raj Prabhu, Maria Tsfasman, Catharine Oertel, Timo Gerkmann, Nale Lehmann-Willenbrock

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07930 2024-10-15 cs.MM cs.CV cs.LG cs.SD eess.AS 78%

Improving Multimodal Learning with Multi-Loss Gradient Modulation

Konstantinos Kontras, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00010 2024-10-02 eess.SP cs.LG 78%

PHemoNet: A Multimodal Network for Physiological Signals

Eleonora Lopez, Aurelio Uncini, Danilo Comminiello

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments The paper has been accepted at RTSI 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06793 2024-09-25 cs.CR cs.IR cs.LG 78%

Adversarial Attacks to Multi-Modal Models

Zhihao Dou, Xin Hu, Haibo Yang, Zhuqing Liu, Minghong Fang

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments To appear in the ACM Workshop on Large AI Systems and Models with Privacy and Safety Analysis 2024 (LAMPS '24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15033 2024-09-24 cs.HC 78%

Immersed in my Ideas: Using Virtual Reality and Multimodal Interactions to Visualize Users' Ideas and Thoughts

Yunhao Xing, Jerrick Ban, Timothy D. Hubbard, Michael Villano, Diego Gomez-Zara

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 24 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07313 2024-08-15 cs.HC 78%

Exploring Large-Scale Language Models to Evaluate EEG-Based Multimodal Data for Mental Health

Yongquan Hu, Shuning Zhang, Ting Dang, Hong Jia, Flora D. Salim, Wen Hu, Aaron J. Quigley

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 6 pages; UbiComp Companion '24, Companion of the 2024 ACM International Joint Conference on Pervasive and Ubiquitous Computing, October 5--9, 2024}{Melbourne, VIC, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14116 2024-08-15 cs.RO cs.HC cs.LG 78%

Learning Multimodal Confidence for Intention Recognition in Human-Robot Interaction

Xiyuan Zhao, Huijun Li, Tianyuan Miao, Xianyi Zhu, Zhikai Wei, Aiguo Song

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15507 2024-06-21 cs.HC 78%

"May I Speak?": Multi-modal Attention Guidance in Social VR Group Conversations

Geonsun Lee, Dae Yeol Lee, Guan-Ming Su, Dinesh Manocha

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00841 2024-06-11 cs.RO 78%

Advantages of Multimodal versus Verbal-Only Robot-to-Human Communication with an Anthropomorphic Robotic Mock Driver

Tim Schreiter, Lucas Morillo-Mendez, Ravi T. Chadalavada, Andrey Rudenko, Erik Billing, Martin Magnusson, Kai O. Arras, Achim J. Lilienthal

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments This paper has been accepted to the 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), which will be held in Busan, South Korea on August 28-31, 2023. For more information, please visit: https://ro-man2023.org/main

Journal ref 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19941 2024-05-31 cs.HC cs.CY 78%

Synthetic Patients: Simulating Difficult Conversations with Multimodal Generative AI for Medical Education

Simon N. Chu, Alex J. Goodell

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12353 2024-05-22 cs.LG 78%

TinyM$^2$Net-V3: Memory-Aware Compressed Multimodal Deep Neural Networks for Sustainable Edge Deployment

Hasib-Al Rashid, Tinoosh Mohsenin

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments Accepted at AAAI 2024 Workshop SAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15174 2024-04-12 cs.RO cs.HC 78%

LaMI: Large Language Models for Multi-Modal Human-Robot Interaction

Chao Wang, Stephan Hasler, Daniel Tanneberg, Felix Ocker, Frank Joublin, Antonello Ceravola, Joerg Deigmoeller, Michael Gienger

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03498 2024-04-05 cs.RO cs.HC 78%

Integrating Large Language Models with Multimodal Virtual Reality Interfaces to Support Collaborative Human-Robot Construction Work

Somin Park, Carol C. Menassa, Vineet R. Kamat

专题命中 音频语音多模态 :multimodal(title,abstract)

Comments 39 pages, 16 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19841 2024-04-01 cs.IR 78%

Dealing with Missing Modalities in Multimodal Recommendation: a Feature Propagation-based Approach

Daniele Malitesta, Emanuele Rossi, Claudio Pomo, Fragkiskos D. Malliaros, Tommaso Di Noia

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19095 2024-03-29 cs.CY 78%

Purposeful remixing with generative AI: Constructing designer voice in multimodal composing

Xiao Tan, Wei Xu, Chaoran Wang

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12609 2024-03-20 cs.LG 78%

SUN Team's Contribution to ABAW 2024 Competition: Audio-visual Valence-Arousal Estimation and Expression Recognition

Denis Dresvyanskiy, Maxim Markitantov, Jiawei Yu, Peitong Li, Heysem Kaya, Alexey Karpov

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract)

Comments 9 pages,

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12267 2024-03-14 cs.RO cs.HC cs.LG 78%

Continuous ErrP detections during multimodal human-robot interaction

Su Kyoung Kim, Michael Maurus, Mathias Trampler, Marc Tabie, Elsa Andrea Kirchner

专题命中 音频语音多模态 :multimodal(title,abstract)

Journal ref Int. Conf. Human-Computer Interaction (2023) 92-101

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07257 2024-03-04 cs.RO 78%

The Audio-Visual BatVision Dataset for Research on Sight and Sound

Amandine Brunetto, Sascha Hornauer, Stella X. Yu, Fabien Moutarde

专题命中 音频语音多模态 :audio-visual(title,abstract)

Comments Project page https://amandinebtto.github.io/Batvision-Dataset/ This version contains camera ready paper

Journal ref 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06306 2024-02-12 cs.IT eess.SP math.IT 78%

Multi-Modal Concurrent Transmission

Majid Nasiri Khormuji, Alberto Giuseppe Perotti, Qin Yi, Branislav Popovic

专题命中 音频语音多模态 :multi-modal(title,abstract)

Comments 6 pages, 4 figures, 1 table

Journal ref 2024 IEEE Wireless Communications and Networking Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10179 2023-12-19 cs.LG 78%

3FM: Multi-modal Meta-learning for Federated Tasks

Minh Tran, Roochi Shah, Zejun Gong

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09740 2023-12-18 cs.RO 78%

VITA: A Multi-modal LLM-based System for Longitudinal, Autonomous, and Adaptive Robotic Mental Well-being Coaching

Micol Spitale, Minja Axelsson, Hatice Gunes

专题命中 音频语音多模态 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.10170 2023-11-20 cs.LG 78%

Improving Unimodal Inference with Multimodal Transformers

Kateryna Chumachenko, Alexandros Iosifidis, Moncef Gabbouj

专题命中 音频语音多模态 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏