arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2505.05441 2025-05-09 cs.HC 50%

GesPrompt: Leveraging Co-Speech Gestures to Augment LLM-Based Interaction in Virtual Reality

Xiyun Hu, Dizhi Ma, Fengming He, Zhengzhe Zhu, Shao-Kang Hsia, Chenfei Zhu, Ziyi Liu, Karthik Ramani

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00153 2025-05-02 cs.HC cs.DC 50%

Audo-Sight: Enabling Ambient Interaction For Blind And Visually Impaired Individuals

Bhanuja Ainary

专题命中 音频语音多模态 :multimodal(abstract)

Comments This thesis was conducted under the guidance of Mohsen Amini Salehi. Special thanks to Minseo Kim and Jacob Bradshaw for their valuable contributions and support throughout the research process. 60 pages, 13 Figures, 2 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09435 2025-04-15 cs.HC 50%

Design Probes for AI-Driven AAC: Addressing Complex Communication Needs in Aphasia

Lei Mao, Jong Ho Lee, Yasmeen Faroqi Shah, Stephanie Valencia

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06379 2025-04-10 cs.HC cs.SY eess.SY 50%

User-Centered Insights into Assistive Navigation Technologies for Individuals with Visual Impairment

Iman Soltani, Johnaton Schofield, Mehran Madani, Daniel Kish, Parisa Emami-Naeini

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04388 2025-04-08 cs.CR 50%

Who's Watching You Zoom? Investigating Privacy of Third-Party Zoom Apps

Saharsh Goenka, Adit Prabhu, Payge Sakurai, Mrinaal Ramachandran, Rakibul Hasan

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01708 2025-04-03 cs.RO cs.HC cs.LG 50%

TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication

Petr Vanc, Karla Stepanova

机构 * Czech Technical University in Prague(布拉格捷克技术大学) Czech Institute of Informatics, Robotics, and Cybernetics(捷克信息学、机器人学与控制论研究所)

专题命中 音频语音多模态 :multimodal(abstract)

Comments 8 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19398 2025-03-26 cs.HC 50%

CyanKitten: AI-Driven Markerless Motion Capture for Improved Elderly Well-Being

Mengyao Guo, Yu Nie, Jinda Han, Zongxing Li, Ze Gao

专题命中 音频语音多模态 :multimodal(abstract)

Comments Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, April 26-May 1, 2025, Yokohama, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.03372 2025-03-25 cs.LG 50%

TartanAviation: Image, Speech, and ADS-B Trajectory Datasets for Terminal Airspace Operations

Jay Patrikar, Joao Dantas, Brady Moon, Milad Hamidi, Sourish Ghosh, Nikhil Keetha, Ian Higgins, Atharva Chandak, Takashi Yoneyama, Sebastian Scherer

机构 * Carnegie Mellon University(卡内基梅隆大学) Aeronautics Institute of Technology(航空技术学院)

专题命中 音频语音多模态 :multi-modal(abstract)

Comments 8 pages, 6 figures, 2 tables

Journal ref Scientific Data volume 12, Article number: 468 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09150 2025-03-13 cs.HC 50%

AdaptAI: A Personalized Solution to Sense Your Stress, Fix Your Mess, and Boost Productivity

Rushiraj Gadhvi, Soham Petkar, Priyansh Desai, Shreyas Ramachandran, Siddharth Siddharth

专题命中 音频语音多模态 :multimodal(abstract)

Comments Accepted for publication at the ACM Conference on Human Factors in Computing Systems (CHI) Late Breaking Work 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06179 2025-02-11 cs.HC 50%

Actual Achieved Gain and Optimal Perceived Gain: Modeling Human Take-over Decisions Towards Automated Vehicles' Suggestions

Shuning Zhang, Xin Yi, Shixuan Li, Chuye Hong, Gujun Chen, Jiarui Liu, Xueyang Wang, Yongquan Hu, Yuntao Wang, Hewu Li

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03965 2025-02-07 cs.LG 50%

Innovative Framework for Early Estimation of Mental Disorder Scores to Enable Timely Interventions

Himanshi Singh, Sadhana Tiwari, Sonali Agarwal, Ritesh Chandra, Sanjay Kumar Sonbhadra, Vrijendra Singh

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09930 2025-02-05 cs.HC 50%

TeamVision: An AI-powered Learning Analytics System for Supporting Reflection in Team-based Healthcare Simulation

Vanessa Echeverria, Linxuan Zhao, Riordan Alfredo, Mikaela Milesi, Yuequiao Jin, Sophie Abel, Jie Fan, Lixiang Yan, Xinyu Li, Samantha Dix, Rosie Wotherspoon, Hollie Jaggard, Abra Osborne, Simon Buckingham Shum, Dragan Gasevic, Roberto Martinez-Maldonado

专题命中 音频语音多模态 :multimodal(abstract)

Comments Accepted to CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12610 2025-02-05 econ.GN cs.CY q-fin.EC 50%

Gender Bias and Property Taxes

Gordon Burtch, Alejandro Zentner

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07051 2025-01-14 cs.RO cs.HC 50%

ROSAnnotator: A Web Application for ROSBag Data Analysis in Human-Robot Interaction

Yan Zhang, Haoqi Li, Ramtin Tabatabaei, Wafa Johal

机构 * University of Melbourne(墨尔本大学)

专题命中 音频语音多模态 :multimodal(abstract)

Comments Accepted to HRI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05496 2025-01-03 cs.CR cs.LG 50%

Implicit Steganography Beyond the Constraints of Modality

Sojeong Song, Seoyun Yang, Chang D. Yoo, Junmo Kim

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院(KAIST))

专题命中 音频语音多模态 :cross-modal(abstract)

Comments 25 pages, Accepted at European Conference on Computer Vision (ECCV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19209 2024-12-30 cs.LG 50%

Context-Aware Deep Learning for Multi Modal Depression Detection

Genevieve Lam, Huang Dongyan, Weisi Lin

专题命中 音频语音多模态 :multi-modal(abstract)

Comments Presented as an Oral at International Conference on Acoustics, Speech and Signal Processing 2019, United Kingdom

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13170 2024-12-18 cs.SI cs.IR 50%

Re-calibrating methodologies in social media research: Challenge the visual, work with Speech

Hongrui Jin

专题命中 音频语音多模态 :multimodal(abstract)

Comments 11 pages (excluding references), 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11036 2024-12-17 math.OC 50%

Stochastic Approximation and Brownian Repulsion based Evolutionary Search

Rajdeep Dutta, T Venkatesh Varma, Saikat Sarkar, Mariya Mamajiwala, Noor Awad, Senthilnath Jayavelu, Debasish Roy

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08223 2024-12-12 cs.HC 50%

Zeitgebers-Based User Time Perception Analysis and Data-Driven Modeling via Transformer in VR

Yi Li, Zengyu Liu, Xiandi Zhu, Ning Xie

专题命中 音频语音多模态 :multimodal(abstract)

Comments 12pages,7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00514 2024-12-03 cs.HC 50%

Alexa, I Wanna See You: Envisioning Smart Home Assistants for the Deaf and Hard-of-Hearing

Tyrone Justin Sta. Maria, Jordan Aiko Deja

专题命中 音频语音多模态 :multimodal(abstract)

Comments 6 pages, 1 figure, 21 references

Journal ref Proceedings of CHIRP 2024: Transforming HCI Research in the Philippines Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18027 2024-11-28 cs.CR 50%

Privacy-preserving Robotic-based Multi-factor Authentication Scheme for Secure Automated Delivery System

Yang Yang, Aryan Mohammadi Pasikhani, Prosanta Gope, Biplab Sikdar

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08343 2024-11-14 q-bio.NC 50%

Brain Treebank: Large-scale intracranial recordings from naturalistic language stimuli

Christopher Wang, Adam Uri Yaari, Aaditya K Singh, Vighnesh Subramaniam, Dana Rosenfarb, Jan DeWitt, Pranav Misra, Joseph R. Madsen, Scellig Stone, Gabriel Kreiman, Boris Katz, Ignacio Cases, Andrei Barbu

专题命中 音频语音多模态 :multimodal(abstract)

Comments 36 pages, 17 figures; Accepted at NeurIPS Dataset and Benchmarks 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.12800 2024-11-12 cs.NI cs.LG 50%

Deep Learning in Physical Layer: Review on Data Driven End-to-End Communication Systems and their Enabling Semantic Applications

Nazmul Islam, Seokjoo Shin

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref IEEE Open Journal of the Communications Society, vol. 5, pp. 4207-4240, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05154 2024-11-11 cs.HC 50%

TelEdge: Haptic Tele-Communication of a Smartphone by Electro-Tactile Stimulation Through the Edges

Taiki Takami, Izumi Mizoguchi, Hiroyuki Kajimoto

专题命中 音频语音多模态 :audio-visual(abstract)

Comments Part of proceedings of 6th International Conference AsiaHaptics 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20545 2024-10-29 cs.HC 50%

ChartA11y: Designing Accessible Touch Experiences of Visualizations with Blind Smartphone Users

Zhuohao Jerry Zhang, John R. Thompson, Aditi Shah, Manish Agrawal, Alper Sarikaya, Jacob O. Wobbrock, Edward Cutrell, Bongshin Lee

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12506 2024-09-20 physics.class-ph 50%

Methodology for 3D sound synthesis of directional acoustic sources by higher-order ambisonics

Philippe Thorner, Eric Bavu, Jean-Baptiste Doc, Christophe Langrenne

专题命中 音频语音多模态 :multimodal(abstract)

Comments Internoise 2024, Aug 2024, Nantes (France), France

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00784 2024-09-04 cs.HC 50%

SonoHaptics: An Audio-Haptic Cursor for Gaze-Based Object Selection in XR

Hyunsung Cho, Naveen Sendhilnathan, Michael Nebeling, Tianyi Wang, Purnima Padmanabhan, Jonathan Browder, David Lindlbauer, Tanya R. Jonker, Kashyap Todi

专题命中 音频语音多模态 :cross-modal(abstract)

Comments UIST 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.17041 2024-09-02 cs.RO 50%

Generative Modeling Perspective for Control and Reasoning in Robotics

Takuma Yoneda

专题命中 音频语音多模态 :multimodal(abstract)

Comments arXiv admin note: text overlap with arXiv:2302.12244

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12124 2024-08-23 cs.LG cs.HC eess.SP 50%

Recording Brain Activity While Listening to Music Using Wearable EEG Devices Combined with Bidirectional Long Short-Term Memory Networks

Jingyi Wang, Zhiqun Wang, Guiran Liu

专题命中 音频语音多模态 :multimodal(abstract)

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.13192 2024-08-19 cs.RO 50%

Anytime, Anywhere: Human Arm Pose from Smartwatch Data for Ubiquitous Robot Control and Teleoperation

Fabian C Weigend, Shubham Sonawani, Michael Drolet, Heni Ben Amor

专题命中 音频语音多模态 :multimodal(abstract)

Comments 8 pages, 10, figures, 1 table, conference: IROS

详情

展开后加载摘要…

URL PDF HTML 收藏