arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2408.03958 2024-08-14 cs.HC cs.LG 50%

Optimizing Emotion Recognition with Wearable Sensor Data: Unveiling Patterns in Body Movements and Heart Rate through Random Forest Hyperparameter Tuning

Zikri Kholifah Nur, Rifki Wijaya, Gia Septiana Wulandari

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 12 pages. Accepted by Jurnal Media Informatika Budidarma (Open Access)

Journal ref Jurnal Media Informatika Budidarma, Vol. 8, No. 3, pp. 1472, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12480 2024-08-08 cs.HC cs.IR 50%

Towards Detecting and Mitigating Cognitive Bias in Spoken Conversational Search

Kaixin Ji, Sachin Pathiyan Cherumanal, Johanne R. Trippas, Danula Hettiachchi, Flora D. Salim, Falk Scholer, Damiano Spina

专题命中 音频语音多模态 :multimodal(abstract)

Comments Extended version of MobileHCI'24 LBW paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09147 2024-07-15 cs.HC cs.IR 50%

AI-Powered Immersive Assistance for Interactive Task Execution in Industrial Environments

Tomislav Duricic, Peter Müllner, Nicole Weidinger, Neven ElSayed, Dominik Kowald, Eduardo Veas

专题命中 音频语音多模态 :multimodal(abstract)

Comments 3 pages, 2 figures, Demo Paper accepted at the 50th European Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.18721 2024-06-05 cs.RO 50%

PhysicsAssistant: An LLM-Powered Interactive Learning Robot for Physics Lab Investigations

Ehsan Latif, Ramviyas Parasuraman, Xiaoming Zhai

专题命中 音频语音多模态 :multimodal(abstract)

Comments Accepted to IEEE RO-MAN Special Session

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17250 2024-05-28 cs.RO cs.SY eess.SY 50%

"Pass the butter": A study on desktop-classic multitasking robotic arm based on advanced YOLOv7 and BERT

Haohua Que, Wenbin Pan, Jie Xu, Hao Luo, Pei Wang, Li Zhang

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.09041 2024-04-16 cs.CR cs.HC 50%

Exploiting Out-of-band Motion Sensor Data to De-anonymize Virtual Reality Users

Mohd Sabra, Nisha Vinayaga Sureshkanth, Ari Sharma, Anindya Maiti, Murtuza Jadliwala

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08623 2024-04-15 cs.HC 50%

Mixing Modes: Active and Passive Integration of Speech, Text, and Visualization for Communicating Data Uncertainty

Chase Stokes, Chelsea Sanker, Bridget Cogley, Vidya Setlur

专题命中 音频语音多模态 :multimodal(abstract)

Comments 5 pages, 2 figures, accepted at Eurographics Conference on Visualization (EuroVis) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02317 2024-04-15 cs.HC 50%

From Delays to Densities: Exploring Data Uncertainty through Speech, Text, and Visualization

Chase Stokes, Chelsea Sanker, Bridget Cogley, Vidya Setlur

专题命中 音频语音多模态 :multimodal(abstract)

Comments 14 pages, 3 figures, accepted at Eurographics Conference on Visualization (EuroVis) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08626 2024-03-14 physics.med-ph 50%

Pulse-echo ultrasound attenuation tomography

Naiara Korta Martiartu, Parisa Salemi Yolgunlu, Martin Frenz, Michael Jaeger

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02274 2024-03-05 cs.RO cs.LG 50%

NatSGD: A Dataset with Speech, Gestures, and Demonstrations for Robot Learning in Natural Human-Robot Interaction

Snehesh Shrestha, Yantian Zha, Saketh Banagiri, Ge Gao, Yiannis Aloimonos, Cornelia Fermuller

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01615 2024-03-05 cs.LG cs.DC 50%

Partial Federated Learning

Tiantian Feng, Anil Ramakrishna, Jimit Majmudar, Charith Peris, Jixuan Wang, Clement Chung, Richard Zemel, Morteza Ziyadi, Rahul Gupta

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17862 2024-02-07 cs.HC 50%

Is Silent eHMI Enough? A Passenger-Centric Study on Effective eHMI for Autonomous Personal Mobility Vehicles in the Field

Hailong Liu, Yang Li, Zhe Zeng, Hao Cheng, Chen Peng, Takahiro Wada

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15201 2024-01-30 cs.HC 50%

Automatically Detecting Confusion and Conflict During Collaborative Learning Using Linguistic, Prosodic, and Facial Cues

Yingbo Ma, Yukyeong Song, Mehmet Celepkolu, Kristy Elizabeth Boyer, Eric Wiebe, Collin F. Lynch, Maya Israel

专题命中 音频语音多模态 :multimodal(abstract)

Comments 27 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15194 2024-01-30 cs.HC 50%

Multimodality in Group Communication Research

Robin Lange, Brooke Foucault Welles, Gyanendra Sharma, Richard J. Radke, Javier O. Garcia, Christoph Riedl

专题命中 音频语音多模态 :multimodal(abstract)

Comments 27 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.00224 2024-01-30 cs.HC 50%

"Hey Model!" - Natural User Interactions and Agency in Accessible Interactive 3D Models

Samuel Reinders, Matthew Butler, Kim Marriott

专题命中 音频语音多模态 :multi-modal(abstract)

Comments Paper presented at ACM CHI 2020: Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, ACM, New York, April 2020; Replacement: typos corrected, character encoding

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18285 2023-12-01 cs.RO 50%

Co-speech gestures for human-robot collaboration

A. Ekrekli, A. Angleraud, G. Sharma, R. Pieters

专题命中 音频语音多模态 :multi-modal(abstract)

Comments 5 pages, accepted to IEEE International Conference on Robotics Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15780 2023-11-28 cs.RO cs.SE cs.SY eess.IV eess.SY 50%

Modular Customizable ROS-Based Framework for Rapid Development of Social Robots

Mahta Akhyani, Hadi Moradi

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15752 2023-11-28 eess.SP 50%

Insights into Age-Related Functional Brain Changes during Audiovisual Integration Tasks: A Comprehensive EEG Source-Based Analysis

Prerna Singh, Ayush Tripathi, Lalan Kumar, Tapan Kumar Gandhi

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01052 2023-11-17 stat.ML cs.LG 50%

Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysis

Victor Letzelter, Mathieu Fontaine, Mickaël Chen, Patrick Pérez, Slim Essid, Gaël Richard

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref Advances in neural information processing systems, Dec 2023, New Orleans, United States

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02819 2023-11-07 cs.CE 50%

An Exploration of Multimodality and Data Augmentation for Dementia Classification

Kaiying Lin, Peter Washington

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02400 2023-11-07 cs.CY 50%

From Plate to Production: Artificial Intelligence in Modern Consumer-Driven Food Systems

Weiqing Min, Pengfei Zhou, Leyi Xu, Tao Liu, Tianhao Li, Mingyu Huang, Ying Jin, Yifan Yi, Min Wen, Shuqiang Jiang, Ramesh Jain

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16956 2023-10-27 cs.CY cs.DB 50%

Datastore Design for Analysis of Police Broadcast Audio at Scale

Ayah Ahmad, Christopher Graziul, Margaret Beale Spencer

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16677 2023-10-26 cs.HC 50%

Machine Learning Approaches for Fine-Grained Symptom Estimation in Schizophrenia: A Comprehensive Review

Niki Maria Foteinopoulou, Ioannis Patras

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 19 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07005 2023-10-12 cs.CR cs.LG 50%

Sound-skwatter (Did You Mean: Sound-squatter?) AI-powered Generator for Phishing Prevention

Rodolfo Valentim, Idilio Drago, Marco Mellia, Federico Cerutti

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09781 2023-09-19 cs.HC 50%

Talking to Data Visualizations: Opportunities and Challenges

Gabriela Molina León, Petra Isenberg, Andreas Breiter

专题命中 音频语音多模态 :multimodal(abstract)

Comments Accepted at the MERCADO workshop at IEEE VIS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09646 2023-09-19 cs.RO 50%

Concurrent Haptic, Audio, and Visual Data Set During Bare Finger Interaction with Textured Surfaces

Alexis W. M. Devillard, Aruna Ramasamy, Damien Faux, Vincent Hayward, Etienne Burdet

专题命中 音频语音多模态 :multi-modal(abstract)

Journal ref 2023 IEEE World Haptics Conference (WHC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.02451 2023-08-23 cs.LG 50%

Tensorized LSSVMs for Multitask Regression

Jiani Liu, Qinghua Tao, Ce Zhu, Yipeng Liu, Johan A. K. Suykens

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 2023, pp. 1-5

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06246 2023-08-14 cs.HC 50%

ARGUS: Visualization of AI-Assisted Task Guidance in AR

Sonia Castelo, Joao Rulff, Erin McGowan, Bea Steers, Guande Wu, Shaoyu Chen, Iran Roman, Roque Lopez, Ethan Brewer, Chen Zhao, Jing Qian, Kyunghyun Cho, He He, Qi Sun, Huy Vo, Juan Bello, Michael Krone, Claudio Silva

专题命中 音频语音多模态 :multimodal(abstract)

Comments 11 pages, 8 figures. This is the author's version of the article of the article that has been accepted for publication in IEEE Transactions on Visualization and Computer Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14718 2023-07-28 cs.HC 50%

Towards a New Interface for Music Listening: A User Experience Study on YouTube

Ahyeon Choi, Eunsik Shin, Haesun Joung, Joongseek Lee, Kyogu Lee

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 6 pages without reference, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07723 2023-07-18 cs.HC 50%

Screen or No Screen? Lessons Learnt from a Real-World Deployment Study of Using Voice Assistants With and Without Touchscreen for Older Adults

Chen Chen, Ella T. Lifset, Yichen Han, Arkajyoti Roy, Michael Hogarth, Alison A. Moore, Emilia Farcas, Nadir Weibel

专题命中 音频语音多模态 :multimodal(abstract)

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏