arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4600 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4600 篇

2405.17250 2024-05-28 cs.RO cs.SY eess.SY 50%

"Pass the butter": A study on desktop-classic multitasking robotic arm based on advanced YOLOv7 and BERT

Haohua Que, Wenbin Pan, Jie Xu, Hao Luo, Pei Wang, Li Zhang

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.09041 2024-04-16 cs.CR cs.HC 50%

Exploiting Out-of-band Motion Sensor Data to De-anonymize Virtual Reality Users

Mohd Sabra, Nisha Vinayaga Sureshkanth, Ari Sharma, Anindya Maiti, Murtuza Jadliwala

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08623 2024-04-15 cs.HC 50%

Mixing Modes: Active and Passive Integration of Speech, Text, and Visualization for Communicating Data Uncertainty

Chase Stokes, Chelsea Sanker, Bridget Cogley, Vidya Setlur

专题命中 音频语音多模态 :multimodal(abstract)

Comments 5 pages, 2 figures, accepted at Eurographics Conference on Visualization (EuroVis) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02317 2024-04-15 cs.HC 50%

From Delays to Densities: Exploring Data Uncertainty through Speech, Text, and Visualization

Chase Stokes, Chelsea Sanker, Bridget Cogley, Vidya Setlur

专题命中 音频语音多模态 :multimodal(abstract)

Comments 14 pages, 3 figures, accepted at Eurographics Conference on Visualization (EuroVis) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08626 2024-03-14 physics.med-ph 50%

Pulse-echo ultrasound attenuation tomography

Naiara Korta Martiartu, Parisa Salemi Yolgunlu, Martin Frenz, Michael Jaeger

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02274 2024-03-05 cs.RO cs.LG 50%

NatSGD: A Dataset with Speech, Gestures, and Demonstrations for Robot Learning in Natural Human-Robot Interaction

Snehesh Shrestha, Yantian Zha, Saketh Banagiri, Ge Gao, Yiannis Aloimonos, Cornelia Fermuller

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01615 2024-03-05 cs.LG cs.DC 50%

Partial Federated Learning

Tiantian Feng, Anil Ramakrishna, Jimit Majmudar, Charith Peris, Jixuan Wang, Clement Chung, Richard Zemel, Morteza Ziyadi, Rahul Gupta

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17862 2024-02-07 cs.HC 50%

Is Silent eHMI Enough? A Passenger-Centric Study on Effective eHMI for Autonomous Personal Mobility Vehicles in the Field

Hailong Liu, Yang Li, Zhe Zeng, Hao Cheng, Chen Peng, Takahiro Wada

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15201 2024-01-30 cs.HC 50%

Automatically Detecting Confusion and Conflict During Collaborative Learning Using Linguistic, Prosodic, and Facial Cues

Yingbo Ma, Yukyeong Song, Mehmet Celepkolu, Kristy Elizabeth Boyer, Eric Wiebe, Collin F. Lynch, Maya Israel

专题命中 音频语音多模态 :multimodal(abstract)

Comments 27 pages, 7 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.15194 2024-01-30 cs.HC 50%

Multimodality in Group Communication Research

Robin Lange, Brooke Foucault Welles, Gyanendra Sharma, Richard J. Radke, Javier O. Garcia, Christoph Riedl

专题命中 音频语音多模态 :multimodal(abstract)

Comments 27 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.00224 2024-01-30 cs.HC 50%

"Hey Model!" - Natural User Interactions and Agency in Accessible Interactive 3D Models

Samuel Reinders, Matthew Butler, Kim Marriott

专题命中 音频语音多模态 :multi-modal(abstract)

Comments Paper presented at ACM CHI 2020: Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, ACM, New York, April 2020; Replacement: typos corrected, character encoding

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18285 2023-12-01 cs.RO 50%

Co-speech gestures for human-robot collaboration

A. Ekrekli, A. Angleraud, G. Sharma, R. Pieters

专题命中 音频语音多模态 :multi-modal(abstract)

Comments 5 pages, accepted to IEEE International Conference on Robotics Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15780 2023-11-28 cs.RO cs.SE cs.SY eess.IV eess.SY 50%

Modular Customizable ROS-Based Framework for Rapid Development of Social Robots

Mahta Akhyani, Hadi Moradi

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15752 2023-11-28 eess.SP 50%

Insights into Age-Related Functional Brain Changes during Audiovisual Integration Tasks: A Comprehensive EEG Source-Based Analysis

Prerna Singh, Ayush Tripathi, Lalan Kumar, Tapan Kumar Gandhi

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01052 2023-11-17 stat.ML cs.LG 50%

Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysis

Victor Letzelter, Mathieu Fontaine, Mickaël Chen, Patrick Pérez, Slim Essid, Gaël Richard

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref Advances in neural information processing systems, Dec 2023, New Orleans, United States

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02819 2023-11-07 cs.CE 50%

An Exploration of Multimodality and Data Augmentation for Dementia Classification

Kaiying Lin, Peter Washington

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02400 2023-11-07 cs.CY 50%

From Plate to Production: Artificial Intelligence in Modern Consumer-Driven Food Systems

Weiqing Min, Pengfei Zhou, Leyi Xu, Tao Liu, Tianhao Li, Mingyu Huang, Ying Jin, Yifan Yi, Min Wen, Shuqiang Jiang, Ramesh Jain

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16956 2023-10-27 cs.CY cs.DB 50%

Datastore Design for Analysis of Police Broadcast Audio at Scale

Ayah Ahmad, Christopher Graziul, Margaret Beale Spencer

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16677 2023-10-26 cs.HC 50%

Machine Learning Approaches for Fine-Grained Symptom Estimation in Schizophrenia: A Comprehensive Review

Niki Maria Foteinopoulou, Ioannis Patras

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 19 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07005 2023-10-12 cs.CR cs.LG 50%

Sound-skwatter (Did You Mean: Sound-squatter?) AI-powered Generator for Phishing Prevention

Rodolfo Valentim, Idilio Drago, Marco Mellia, Federico Cerutti

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09781 2023-09-19 cs.HC 50%

Talking to Data Visualizations: Opportunities and Challenges

Gabriela Molina León, Petra Isenberg, Andreas Breiter

专题命中 音频语音多模态 :multimodal(abstract)

Comments Accepted at the MERCADO workshop at IEEE VIS 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.09646 2023-09-19 cs.RO 50%

Concurrent Haptic, Audio, and Visual Data Set During Bare Finger Interaction with Textured Surfaces

Alexis W. M. Devillard, Aruna Ramasamy, Damien Faux, Vincent Hayward, Etienne Burdet

专题命中 音频语音多模态 :multi-modal(abstract)

Journal ref 2023 IEEE World Haptics Conference (WHC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.02451 2023-08-23 cs.LG 50%

Tensorized LSSVMs for Multitask Regression

Jiani Liu, Qinghua Tao, Ce Zhu, Yipeng Liu, Johan A. K. Suykens

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 2023, pp. 1-5

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06246 2023-08-14 cs.HC 50%

ARGUS: Visualization of AI-Assisted Task Guidance in AR

Sonia Castelo, Joao Rulff, Erin McGowan, Bea Steers, Guande Wu, Shaoyu Chen, Iran Roman, Roque Lopez, Ethan Brewer, Chen Zhao, Jing Qian, Kyunghyun Cho, He He, Qi Sun, Huy Vo, Juan Bello, Michael Krone, Claudio Silva

专题命中 音频语音多模态 :multimodal(abstract)

Comments 11 pages, 8 figures. This is the author's version of the article of the article that has been accepted for publication in IEEE Transactions on Visualization and Computer Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14718 2023-07-28 cs.HC 50%

Towards a New Interface for Music Listening: A User Experience Study on YouTube

Ahyeon Choi, Eunsik Shin, Haesun Joung, Joongseek Lee, Kyogu Lee

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 6 pages without reference, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07723 2023-07-18 cs.HC 50%

Screen or No Screen? Lessons Learnt from a Real-World Deployment Study of Using Voice Assistants With and Without Touchscreen for Older Adults

Chen Chen, Ella T. Lifset, Yichen Han, Arkajyoti Roy, Michael Hogarth, Alison A. Moore, Emilia Farcas, Nadir Weibel

专题命中 音频语音多模态 :multimodal(abstract)

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.05451 2023-07-12 cs.HC 50%

Detection Threshold of Audio Haptic Asynchrony in a Driving Context

Gyanendra Sharma, Hiroshi Yasuda, Manuel Kuehner

专题命中 音频语音多模态 :multimodal(abstract)

Comments 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.00124 2023-06-22 cs.LO 50%

Axiomatizing Hybrid XPath with Data

Carlos Areces, Raul Fervari

专题命中 音频语音多模态 :multi-modal(abstract)

Journal ref Logical Methods in Computer Science, Volume 17, Issue 3 (July 20, 2021) lmcs:6259

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11600 2023-06-16 cs.SI cs.CR 50%

Creative beyond TikToks: Investigating Adolescents' Social Privacy Management on TikTok

Nico Ebert, Tim Geppert, Joanna Strycharz, Melanie Knieps, Michael Hönig, Elke Brucker-Kley

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17180 2023-05-30 cs.HC 50%

Exploring Human Response Times to Combinations of Audio, Haptic, and Visual Stimuli from a Mobile Device

Kyle T. Yoshida, Joel X. Kiernan, Allison M. Okamura, Cara M. Nunez

专题命中 音频语音多模态 :multi-modal(abstract)

Comments Accepted to World Haptics Conference 2023

详情

展开后加载摘要…

URL PDF HTML 收藏