arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4600 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4600 篇

2105.02139 2021-05-06 cs.HC 50%

Mixing Modalities of 3D Sketching and Speech for Interactive Model Retrieval in Virtual Reality

Daniele Giunchi, Alejandro Sztrajman, Stuart James, Anthony Steed

专题命中 音频语音多模态 :multimodal(abstract)

Comments Published at IMX 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.13836 2021-04-29 cs.HC 50%

The EMPATHIC Project: Building an Expressive, Advanced Virtual Coach to Improve Independent Healthy-Life-Years of the Elderly

Luisa Brinkschulte, Natascha Mariacher, Stephan Schlögl, María Inés Torres, Raquel Justo, Javier Mikel Olaso, Anna Esposito, Gennaro Cordasco, Gérard Chollet, Cornelius Glackin, Colin Pickard, Dijana Petrovska-Delacretaz, Mohamed Amine Hmani, Ayment Mtibaa, Anaïs Fernandez, Daria Kyslitska, Begoña Fernandez-Ruanova, Jofre Tenorio-Laranga, Mari Aksnes, Maria Stylianou Korsnes, Miriam Reiner, Fredrik Lindner, Olivier Deroo, Olga Gordeeva

专题命中 音频语音多模态 :multimodal(abstract)

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.11537 2021-04-05 cs.HC physics.app-ph physics.bio-ph 50%

Reduced Graphene Oxide Tattoo as Wearable Proximity Sensor

Vaishakh Kedambaimoole, Neelotpala Kumar, Vijay Shirhatti, Suresh Nuthalapati, Saurabh Kumar, Mangalore Manjunatha Nayak, Prosenjit Sen, Deji Akinwande, Konandur Rajanna

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.14189 2021-03-29 cs.DB cs.LG 50%

DBATES: DataBase of Audio features, Text, and visual Expressions in competitive debate Speeches

Taylan K. Sen, Gazi Naven, Luke Gerstner, Daryl Bagley, Raiyan Abdul Baten, Wasifur Rahman, Kamrul Hasan, Kurtis G. Haut, Abdullah Mamun, Samiha Samrose, Anne Solbu, R. Eric Barnes, Mark G. Frank, Ehsan Hoque

专题命中 音频语音多模态 :multimodal(abstract)

Comments 12 pages, 5 figures, 4 tables, under-going major revision for TAC

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.10245 2021-01-26 cs.HC 50%

AirWare: Utilizing Embedded Audio and Infrared Signals for In-Air Hand-Gesture Recognition

Nibhrat Lohia, Raunak Mundada, Arya D. McCarthy, Eric C. Larson

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.08319 2021-01-22 cs.HC 50%

Communication Aid for Non-English Speaking Newcomers

Munira Al-Ageili, Malek Mouhoub

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.06283 2021-01-19 cs.HC 50%

Data@Hand: Fostering Visual Exploration of Personal Data on Smartphones Leveraging Speech and Touch Interaction

Young-Ho Kim, Bongshin Lee, Arjun Srinivasan, Eun Kyoung Choe

专题命中 音频语音多模态 :multimodal(abstract)

Comments To appear in ACM CHI 2021 Conference on Human Factors in Computing Systems; 16 pages, 6 figures, 5 tables

Journal ref In CHI Conference on Human Factors in Computing Systems (CHI '21), May 8-13, 2021, Yokohama, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.02252 2021-01-08 cs.RO 50%

Playing with Food: Learning Food Item Representations through Interactive Exploration

Amrita Sawhney, Steven Lee, Kevin Zhang, Manuela Veloso, Oliver Kroemer

专题命中 音频语音多模态 :multimodal(abstract)

Comments 12 pages, 8 figures, 2 tables, to be published in Proceedings of International Symposium on Experimental Robotics (ISER) 2020, project website located here: https://sites.google.com/view/playing-with-food/home

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.10723 2020-11-25 cs.HC 50%

NL4DV: A Toolkit for Generating Analytic Specifications for Data Visualization from Natural Language Queries

Arpit Narechania, Arjun Srinivasan, John Stasko

专题命中 音频语音多模态 :multimodal(abstract)

Comments 11 pages, 10 figures. Proceedings of IEEE VIS'2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.08518 2020-11-09 cs.HC 50%

Emotion-Recognition Using Smart Watch Sensor Data: Mixed-Design Study

Juan C. Quiroz, Elena Geangu, Min Hooi Yong

专题命中 音频语音多模态 :audio-visual(abstract)

Comments Published in JMIR Mental Health

Journal ref JMIR Ment Health (2018);5(3):e10153

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.01688 2020-07-06 cs.CR cs.CY 50%

Online publication of court records: circumventing the privacy-transparency trade-off

Tristan Allard, Louis Béziaud, Sébastien Gambs

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.10293 2020-06-19 cs.LG stat.ML 50%

GAT-GMM: Generative Adversarial Training for Gaussian Mixture Models

Farzan Farnia, William Wang, Subhro Das, Ali Jadbabaie

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.05442 2020-06-11 cs.LG stat.ML 50%

Tensor train decompositions on recurrent networks

Alejandro Murua, Ramchalam Ramakrishnan, Xinlin Li, Rui Heng Yang, Vahid Partovi Nia

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.08578 2020-05-12 cs.CR cs.HC cs.LG 50%

Sensor-based Continuous Authentication of Smartphones' Users Using Behavioral Biometrics: A Contemporary Survey

Mohammed Abuhamad, Ahmed Abusnaina, DaeHun Nyang, David Mohaisen

专题命中 音频语音多模态 :multimodal(abstract)

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.06655 2020-02-18 cs.HC 50%

Concurrent Crossmodal Feedback Assists Target-searching: Displaying Distance Information Through Visual, Auditory and Haptic Modalities

Feng Feng, Tony Stockman

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.08858 2019-09-20 cond-mat.mtrl-sci cond-mat.mes-hall physics.app-ph 50%

Ultimate Photo-Thermo-Acoustic Efficiency of Graphene Aerogels

Francesco De Nicola, Lorenzo Donato Tenuzzo, Ilenia Viola, Rujing Zhang, Hongwei Zhu, Augusto Marcelli, Stefano Lupi

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 6 pages, 4 figures

Journal ref Scientific Reports 9, 13386 (2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.04582 2019-09-12 eess.SP cs.HC 50%

Consumer Grade Brain Sensing for Emotion Recognition

Payongkit Lakhan, Nannapas Banluesombatkul, Vongsagon Changniam, Ratwade Dhithijaiyratn, Pitshaporn Leelaarporn, Ekkarat Boonchieng, Supanida Hompoonsup, Theerawit Wilaiprasitporn

专题命中 音频语音多模态 :audio-visual(abstract)

Journal ref IEEE Sensor Journal, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.03902 2019-09-10 cs.NI eess.SP 50%

Analyzing the Trade-offs in Using Millimeter Wave Directional Links for High Data Rate Tactile Internet Applications

Kishor Chandra Joshi, Solmaz Niknam, R. Venkatesha Prasad, Balasubramaniam Natarajan

专题命中 音频语音多模态 :audio-visual(abstract)

Comments IEEE Transactions on Industrial Informatics, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.01963 2019-08-07 cs.HC physics.ed-ph 50%

Using the Makerspace to Create Educational Open-source Software for Electrical Circuits: A Learning Experience

Dana Conard, Blake Vollmer, Corbin Shatto, Hannah Bowman, Sara Kassis

专题命中 音频语音多模态 :multimodal(abstract)

Comments Presented as poster at International Symposium on Academic Makerspaces(ISAM) 2018 - https://isam2018.hemi-makers.org/papers-posters/

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.01606 2019-05-28 cs.LG 50%

DeepKey: An EEG and Gait Based Dual-Authentication System

Xiang Zhang, Lina Yao, Chaoran Huang, Tao Gu, Zheng Yang, Yunhao Liu

专题命中 音频语音多模态 :multimodal(abstract)

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.01352 2019-05-07 cs.HC 50%

PAL: A Wearable Platform for Real-time, Personalized and Context-Aware Health and Cognition Support

Mina Khan, Glenn Fernandes, Utkarsh Sarawgi, Prudhvi Rampey, Pattie Maes

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.11943 2018-11-30 cs.HC 50%

The haptic paradigm in education: Challenges and case studies

Felix G. Hamza-Lup, Ioana A. Stanescu

专题命中 音频语音多模态 :multi-modal(abstract)

Journal ref Internet and Higher Education Journal (2010), Vol. 13(1), pp. 78-81 (ISSN 1096-7516)

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.01033 2018-11-06 cs.HC 50%

Building an Argument for the Use of Science Fiction in HCI Education

Philipp Jordan, Paula Alexandra Silva

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 6 pages, 1 table, IHSI 2019 accepted submission

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.03724 2018-10-25 cs.FL cs.SE 50%

DReAM: Dynamic Reconfigurable Architecture Modeling (full paper)

Rocco De Nicola, Alessandro Maggi, Joseph Sifakis

专题命中 音频语音多模态 :multi-modal(abstract)

Comments 29 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.10408 2018-09-26 cs.RO cs.HC 50%

A Neurorobotic Experiment for Crossmodal Conflict Resolution in Complex Environments

German I. Parisi, Pablo Barros, Di Fu, Sven Magg, Haiyan Wu, Xun Liu, Stefan Wermter

专题命中 音频语音多模态 :audio-visual(abstract)

Comments Accepted at the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2018), Madrid, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.07276 2018-09-21 cs.IR cs.LG cs.SD stat.ML 50%

Music Mood Detection Based On Audio And Lyrics With Deep Neural Net

Rémi Delbouys, Romain Hennequin, Francesco Piccoli, Jimena Royo-Letelier, Manuel Moussallam

专题命中 音频语音多模态 :multimodal(abstract)

Comments Published in ISMIR 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.04966 2018-09-14 cs.CR 50%

Real-Time Lightweight Chaotic Encryption for 5G IoT Enabled Lip-Reading Driven Secure Hearing-Aid

Ahsan Adeel, Jawad Ahmad, Amir Hussain

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.09060 2018-06-26 cs.LG stat.ML 50%

Disentangled VAE Representations for Multi-Aspect and Missing Data

Samuel K. Ainsworth, Nicholas J. Foti, Emily B. Fox

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.00296 2018-06-24 cs.HC 50%

Dišimo: Anchoring Our Breath

Jelena Mladenovic, Jérémy Frey, Jessica Cauchard

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref CHI '18 Interactivity - SIGCHI Conference on Human Factors in Computing System, Apr 2018, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.10285 2018-04-30 math.LO 50%

Pointwise intersection in neighbourhood modal logic

Frederik Van De Putte, Dominik Klein

专题命中 音频语音多模态 :multi-modal(abstract)

Comments Submitted to Advances in Modal Logic 2018

详情

展开后加载摘要…

URL PDF HTML 收藏