arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

1906.11759 2019-06-28 q-bio.NC cs.IR cs.LG cs.SD eess.AS stat.ML 57%

Low-dimensional Embodied Semantics for Music and Language

Francisco Afonso Raposo, David Martins de Matos, Ricardo Ribeiro

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 6 pages, 1 figure, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.10606 2019-06-26 eess.AS cs.DB cs.LG cs.SD 57%

DALI: a large Dataset of synchronized Audio, LyrIcs and notes, automatically created using teacher-student machine learning paradigm

Gabriel Meseguer-Brocal, Alice Cohen-Hadria, Geoffroy Peeters

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Journal ref Proceedings of the 19th International Society for Music Information Retrieval Conference, ISMIR, Paris, France, pp. 431-437, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.09832 2019-06-25 cs.CL cs.LG cs.SD 57%

A computational model of early language acquisition from audiovisual experiences of young infants

Okko Räsänen, Khazar Khorrami

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.10548 2019-06-06 stat.ML cs.CL cs.LG 57%

Latent Normalizing Flows for Discrete Sequences

Zachary M. Ziegler, Alexander M. Rush

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.01778 2019-06-04 cs.HC cs.AI 57%

Recognition of Advertisement Emotions with Application to Computational Advertising

Abhinav Shukla, Shruti Shriya Gullapuram, Harish Katti, Mohan Kankanhalli, Stefan Winkler, Ramanathan Subramanian

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

Comments Under consideration for publication in IEEE Trans. Affective Computing. arXiv admin note: text overlap with arXiv:1709.01684

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.06860 2019-05-17 eess.AS cs.LG cs.SD stat.ML 57%

Speaker-Independent Speech-Driven Visual Speech Synthesis using Domain-Adapted Acoustic Models

Ahmed Hussen Abdelaziz, Barry-John Theobald, Justin Binder, Gabriele Fanelli, Paul Dixon, Nicholas Apostoloff, Thibaut Weise, Sachin Kajareker

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 9 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.04192 2019-05-13 cs.LG cs.AI cs.SD stat.ML 57%

Do Autonomous Agents Benefit from Hearing?

Abraham Woubie, Anssi Kanervisto, Janne Karttunen, Ville Hautamaki

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.01850 2019-05-10 cs.SD cs.LG eess.AS 57%

End-to-End Sound Source Separation Conditioned On Instrument Labels

Olga Slizovskaia, Leo Kim, Gloria Haro, Emilia Gomez

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 5 pages, 2 figures, 2 tables, ICASSP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.10715 2019-04-25 cs.MM 57%

Système d'indexation et de recherche de vidéo intégrant un système gestuel pour les personnes handicapées

Mohamed Hamroun, Mohamed Salim Bouhlel

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

Comments 6 pages, in French. International conference Human-Machine Interaction and Image IHMIM14, May 2014, Hammamet, Tunisia

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.00202 2019-04-02 cs.SD eess.AS 57%

Static Visual Spatial Priors for DoA Estimation

Pawel Swietojanski, Ondrej Miksik

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments 6 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.07981 2019-03-20 cs.LG cs.SD eess.AS 57%

Twins Recognition with Multi Biometric System: Handcrafted-Deep Learning Based Multi Algorithm with Voice-Ear Recognition Based Multi Modal

Cihan Akın, Umit Kacar, Murvet Kirci

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.12139 2018-11-30 cs.CV 57%

Two-level Attention with Two-stage Multi-task Learning for Facial Emotion Recognition

Xiaohua Wang, Muzi Peng, Lijuan Pan, Min Hu, Chunhua Jin, Fuji Ren

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.02411 2018-11-07 cs.SD eess.AS 57%

An audio-only method for advertisement detection in broadcast television content

António Ramires, Diogo Cocharro, Matthew E. P. Davies

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Journal ref Proc. of RecPad-2017, Amadora, Portugal, pp. 21-22, October, 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.07455 2018-10-18 cs.CL 57%

Exploring Textual and Speech information in Dialogue Act Classification with Speaker Domain Adaptation

Xuanli He, Quan Hung Tran, William Havard, Laurent Besacier, Ingrid Zukerman, Gholamreza Haffari

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 5 pages, 2 figurs

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.06748 2018-10-17 cs.AI cs.HC q-bio.NC 57%

Assessing the Contribution of Semantic Congruency to Multisensory Integration and Conflict Resolution

Di Fu, Pablo Barros, German I. Parisi, Haiyan Wu, Sven Magg, Xun Liu, Stefan Wermter

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

Comments Workshop on Crossmodal Learning for Intelligent Robotics at IROS'18, Madrid, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.05689 2018-09-18 cs.SD cs.LG eess.AS 57%

Attention as a Perspective for Learning Tempo-invariant Audio Queries

Matthias Dorfer, Jan Hajič, Gerhard Widmer

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments The 2018 Joint Workshop on Machine Learning for Music

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.07773 2018-08-24 cs.CV 57%

EmotiW 2018: Audio-Video, Student Engagement and Group-Level Affect Prediction

Abhinav Dhall, Amanjot Kaur, Roland Goecke, Tom Gedeon

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.06250 2018-08-21 cs.CV 57%

Dynamic Temporal Alignment of Speech to Lips

Tavi Halperin, Ariel Ephrat, Shmuel Peleg

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.06841 2018-07-20 cs.SD eess.AS 57%

Music Style Transfer: A Position Paper

Shuqi Dai, Zheng Zhang, Gus G. Xia

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments In Proceeding of International Workshop on Musical Metacreation (MUME), 2018, Salamanca, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.11688 2018-05-31 eess.IV eess.AS 57%

Towards Lipreading Sentences with Active Appearance Models

George Sterpu, Naomi Harte

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Presented at The 14th International Conference on Auditory-Visual Speech Processing (AVSP 2017)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.00705 2018-05-17 cs.AI cs.LG stat.ML 57%

Investigating Audio, Visual, and Text Fusion Methods for End-to-End Automatic Personality Prediction

Onno Kampman, Elham J. Barezi, Dario Bertero, Pascale Fung

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted at ACL2018 short paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.01122 2018-03-06 eess.AS cs.SD 57%

An Ensemble Framework of Voice-Based Emotion Recognition System for Films and TV Programs

Fei Tao, Gang Liu, Qingen Zhao

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.06767 2018-02-13 cs.CV cs.SD 57%

Seeing Through Noise: Visually Driven Speaker Separation and Enhancement

Aviv Gabbay, Ariel Ephrat, Tavi Halperin, Shmuel Peleg

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Supplementary video: https://www.youtube.com/watch?v=qmsyj7vAzoI

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.09056 2018-01-30 cs.CV 57%

A Multi-Biometrics for Twins Identification Based Speech and Ear

Cihan Akin, Umit Kacar, Murvet Kirci

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.07481 2018-01-24 cs.CV 57%

Survey on Emotional Body Gesture Recognition

Fatemeh Noroozi, Ciprian Adrian Corneanu, Dorota Kamińska, Tomasz Sapiński, Sergio Escalera, Gholamreza Anbarjafari

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.05197 2017-12-18 cs.IR cs.LG cs.SD eess.AS q-bio.NC 57%

Towards Deep Modeling of Music Semantics using EEG Regularizers

Francisco Raposo, David Martins de Matos, Ricardo Ribeiro, Suhua Tang, Yi Yu

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.09443 2017-09-28 cs.CL 57%

Prosodic Features from Large Corpora of Child-Directed Speech as Predictors of the Age of Acquisition of Words

Lea Frermann, Michael C. Frank

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1705.08168 2017-08-02 cs.CV cs.LG 57%

Look, Listen and Learn

Relja Arandjelović, Andrew Zisserman

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Appears in: IEEE International Conference on Computer Vision (ICCV) 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.05050 2016-12-16 cs.LG cs.CV 57%

Towards Score Following in Sheet Music Images

Matthias Dorfer, Andreas Arzt, Gerhard Widmer

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Published In Proceedings of the 17th International Society for Music Information Retrieval Conference (2016)

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.10120 2016-12-01 cs.AI cs.HC 57%

Fusion of EEG and Musical Features in Continuous Music-emotion Recognition

Nattapong Thammasan, Ken-ichi Fukui, Masayuki Numao

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments The short version of this paper is accepted to appear as an abstract in the proceedings of AAAI-17 (student abstract and poster program)

详情

展开后加载摘要…

URL PDF HTML 收藏