arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

1809.04966 2018-09-14 cs.CR 50%

Real-Time Lightweight Chaotic Encryption for 5G IoT Enabled Lip-Reading Driven Secure Hearing-Aid

Ahsan Adeel, Jawad Ahmad, Amir Hussain

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.09060 2018-06-26 cs.LG stat.ML 50%

Disentangled VAE Representations for Multi-Aspect and Missing Data

Samuel K. Ainsworth, Nicholas J. Foti, Emily B. Fox

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.00296 2018-06-24 cs.HC 50%

Dišimo: Anchoring Our Breath

Jelena Mladenovic, Jérémy Frey, Jessica Cauchard

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref CHI '18 Interactivity - SIGCHI Conference on Human Factors in Computing System, Apr 2018, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.10285 2018-04-30 math.LO 50%

Pointwise intersection in neighbourhood modal logic

Frederik Van De Putte, Dominik Klein

专题命中 音频语音多模态 :multi-modal(abstract)

Comments Submitted to Advances in Modal Logic 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.08386 2018-04-24 cs.HC 50%

All Reality: Virtual, Augmented, Mixed (X), Mediated (X,Y), and Multimediated Reality

Steve Mann, Tom Furness, Yu Yuan, Jay Iorio, Zixin Wang

专题命中 音频语音多模态 :multimodal(abstract)

Comments 14 pages, 20 figures, expanded version of a much shorter paper submitted to ACM Multimedia 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1804.00779 2018-04-04 cs.LG stat.ML 50%

Neural Autoregressive Flows

Chin-Wei Huang, David Krueger, Alexandre Lacoste, Aaron Courville

专题命中 音频语音多模态 :multimodal(abstract)

Comments 16 pages, 10 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.08383 2018-03-23 cs.HC 50%

Head-up Displays (HUD) in driving

Marcos Maroto, Enrique Caño, Pavel González, Diego Villegas

专题命中 音频语音多模态 :multimodal(abstract)

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.01662 2018-03-06 cs.HC 50%

Continuous Affect Prediction Using Eye Gaze and Speech

Jonny O'Dwyer, Ronan Flynn, Niall Murray

专题命中 音频语音多模态 :audio-visual(abstract)

Comments Accepted paper for the 2017 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)

Journal ref J. O Dwyer, R. Flynn, and N. Murray, Continuous affect prediction using eye gaze and speech, in 2017 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2017, pp. 2001-2007

详情

展开后加载摘要…

URL PDF HTML 收藏
1508.00883 2017-11-17 cs.CY cs.HC 50%

Equality of Participation Online Versus Face to Face: Condensed Analysis of the Community Forum Deliberative Methods Demonstration

Eric Showers, Nathan Tindall, Todd Davies

专题命中 音频语音多模态 :cross-modal(abstract)

Comments 14 pages, 10 tables, to appear in Efthimios Tambouris, Panos Panagiotopoulos, Øystein Sæbø, Konstantinos Tarabanis, Michela Milano, Theresa Pardo, and Maria Wimmer (Editors), Electronic Participation: Proceedings of the 7th IFIP WG 8.5 International Conference, ePart 2015 (Thessaloniki, August 30-September 2), Springer LNCS Vol. 9249, 2015

Journal ref Lecture Notes in Computer Science 9249:53-67, 2015

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.08836 2017-10-25 cs.CY 50%

A Quantitative Study of the Impact of Social Media Reviews on Brand Perception

Neha Joshi

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.00171 2017-10-03 cs.HC 50%

Confirmation detection in human-agent interaction using non-lexical speech cues

Mara Brandt, Britta Wrede, Franz Kummert, Lars Schillingmann

专题命中 音频语音多模态 :multi-modal(abstract)

Comments 6 pages, Symposium on Natural Communication for Human-Robot Collaboration, AAAI Fall Symposium Series 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.09148 2017-09-27 cs.HC 50%

Emotion-Recognition Using Smart Watch Accelerometer Data: Preliminary Findings

Juan C. Quiroz, Min Hooi Yong, Elena Geangu

专题命中 音频语音多模态 :audio-visual(abstract)

Comments Mental Health and Well-being: Sensing and Intervention, UBICOMP 2017 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.09887 2017-08-01 cs.IR cs.SD 50%

Learning Audio - Sheet Music Correspondences for Score Identification and Offline Alignment

Matthias Dorfer, Andreas Arzt, Gerhard Widmer

专题命中 音频语音多模态 :multi-modal(abstract)

Comments In Proceedings of the 18th International Society for Music Information Retrieval Conference (ISMIR 2017)

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.03219 2017-06-02 stat.ML cs.LG 50%

Sharing Hash Codes for Multiple Purposes

Wikor Pronobis, Danny Panknin, Johannes Kirschnick, Vignesh Srinivasan, Wojciech Samek, Volker Markl, Manohar Kaul, Klaus-Robert Mueller, Shinichi Nakajima

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.08051 2017-05-24 cs.LG stat.ML 50%

Wasserstein Learning of Deep Generative Point Process Models

Shuai Xiao, Mehrdad Farajtabar, Xiaojing Ye, Junchi Yan, Le Song, Hongyuan Zha

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1603.07341 2017-05-24 cs.LG cs.NE stat.ML 50%

Acceleration of Deep Neural Network Training with Resistive Cross-Point Devices

Tayfun Gokmen, Yurii Vlasov

专题命中 音频语音多模态 :multimodal(abstract)

Comments 19 pages, 5 figures, 2 tables

Journal ref Front. Neurosci 10, 333 (2016)

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.04797 2017-04-18 cs.RO 50%

Setting Up Pepper For Autonomous Navigation And Personalized Interaction With Users

Vittorio Perera, Tiago Pereira, Jonathan Connell, Manuela Veloso

专题命中 音频语音多模态 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.05076 2016-12-16 cs.SD 50%

Live Score Following on Sheet Music Images

Matthias Dorfer, Andreas Arzt, Sebastian Böck, Amaury Durand, Gerhard Widmer

专题命中 音频语音多模态 :multi-modal(abstract)

Comments 17th International Society for Music Information Retrieval Conference (ISMIR 2016), Late Breaking/Demo Papers, New York, NY

详情

展开后加载摘要…

URL PDF HTML 收藏
1001.3748 2016-09-08 cs.CY 50%

Enhancing Fine Motor Skills of Wards with Special Needs Using Cluster Model of Cognition

T. R. Gopalakrishnan Nair, N. Sowjanya Rao, Ananda Bukkambudhi

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 14 pages, 6 figures

Journal ref International Conference, Team Tech pp 32, 2009

详情

展开后加载摘要…

URL PDF HTML 收藏
1510.04540 2016-08-26 math.AP 50%

Team organization may help swarms of flies to become invisible in closed waveguides

Lucas Chesnel, Sergei A. Nazarov

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1507.08505 2016-08-09 q-bio.NC 50%

Measures of multisensory integration based on dependent probability summation: from spike counts to reaction times

Hans Colonius, Adele Diederich

专题命中 音频语音多模态 :audio-visual(abstract)

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1607.07801 2016-07-27 cs.SD 50%

ABROA : Audio-Based Room-Occupancy Analysis using Gaussian Mixtures and Hidden Markov Models

Rafael Valle

专题命中 音频语音多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.09437 2016-06-02 cs.CY 50%

Fog Data: Enhancing Telehealth Big Data Through Fog Computing

Harishchandra Dubey, Jing Yang, Nick Constant, Amir Mohammad Amiri, Qing Yang, Kunal Makodiya

专题命中 音频语音多模态 :multi-modal(abstract)

Comments 6 pages, 4 figures in ASE BD&SI '15 Proceedings of the ASE BigData & SocialInformatics 2015, ACM, NY

详情

展开后加载摘要…

URL PDF HTML 收藏
1510.06113 2016-03-02 cs.RO cs.DC 50%

Automated Synchronization of Driving Data Using Vibration and Steering Events

Lex Fridman, Daniel E Brown, William Angell, Irman Abdić, Bryan Reimer, Hae Young Noh

专题命中 音频语音多模态 :multi-modal(abstract)

Comments Accepted for Publication in Elsevier Pattern Recognition Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
1010.3806 2015-07-01 cs.PL 50%

A Logical Foundation for Environment Classifiers

Takeshi Tsukada, Atsushi Igarashi

专题命中 音频语音多模态 :multi-modal(abstract)

Journal ref Logical Methods in Computer Science, Volume 6, Issue 4 (December 18, 2010) lmcs:1065

详情

展开后加载摘要…

URL PDF HTML 收藏
1505.02137 2015-05-29 cs.CY cs.LG 50%

Human Social Interaction Modeling Using Temporal Deep Networks

Mohamed R. Amer, Behjat Siddiquie, Amir Tamrakar, David A. Salter, Brian Lande, Darius Mehri, Ajay Divakaran

专题命中 音频语音多模态 :audio-visual(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1501.05068 2015-01-22 physics.data-an physics.med-ph stat.ML 50%

Difficulties applying recent blind source separation techniques to EEG and MEG

Kevin H. Knuth

专题命中 音频语音多模态 :multimodal(abstract)

Comments 14 pages, 5 figures; Knuth K. H. 1998. Difficulties applying recent blind source separation techniques to EEG and MEG. In: G.J. Erickson, J.T. Rychert and C.R. Smith (eds.), Maximum Entropy and Bayesian Methods, Boise 1997, Kluwer, Dordrecht, pp. 209-222

详情

展开后加载摘要…

URL PDF HTML 收藏
1312.2894 2013-12-11 cs.LO 50%

A Proof Procedure for Hybrid Logic with Binders, Transitivity and Relation Hierarchies (extended version)

Marta Cialdea Mayer

专题命中 音频语音多模态 :multi-modal(abstract)

Comments arXiv admin note: text overlap with arXiv:1210.5734

详情

展开后加载摘要…

URL PDF HTML 收藏
1311.5757 2013-11-25 cs.HC 50%

Lemma 4: Haptic Input + Auditory Display = Musical Instrument?

Paul Vickers

专题命中 音频语音多模态 :multimodal(abstract)

Comments in Haptic and Audio Interaction Design: First International Workshop, HAID 2006, Glasgow, UK, August 31 - September 1, 2006. Proceedings (D. McGookin and S. Brewster, eds.), vol. 4129/2006 of Lecture Notes in Computer Science, pp. 56-67, Springer-Verlag, 2006

详情

展开后加载摘要…

URL PDF HTML 收藏
1310.8585 2013-11-01 cs.HC q-bio.QM 50%

Speech animation using electromagnetic articulography as motion capture data

Ingmar Steiner, Korin Richmond, Slim Ouni

专题命中 音频语音多模态 :multimodal(abstract)

Journal ref AVSP - 12th International Conference on Auditory-Visual Speech Processing - 2013 (2013) 55-60

详情

展开后加载摘要…

URL PDF HTML 收藏