arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2008.02070 2020-08-06 eess.AS cs.LG cs.SD 57%

Content based singing voice source separation via strong conditioning using aligned phonemes

Gabriel Meseguer-Brocal, Geoffroy Peeters

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 21st International Society for Music Information Retrieval Conference 11-15 October 2020, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.03889 2020-08-03 eess.AS cs.SD 57%

Neural Spatio-Temporal Beamformer for Target Speech Separation

Yong Xu, Meng Yu, Shi-Xiong Zhang, Lianwu Chen, Chao Weng, Jianming Liu, Dong Yu

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments accepted to Interspeech2020, Demo: https://yongxuustc.github.io/mtmvdr/

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.08874 2020-07-24 eess.AS cs.LG cs.SD 57%

Speech Emotion Recognition with Dual-Sequence LSTM Architecture

Jianyou Wang, Michael Xue, Ryan Culhane, Enmao Diao, Jie Ding, Vahid Tarokh

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by ICASSP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.06355 2020-07-15 cs.CV 57%

Multiple Sound Sources Localization from Coarse to Fine

Rui Qian, Di Hu, Heinrich Dinkel, Mengyue Wu, Ning Xu, Weiyao Lin

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments to appear in ECCV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.03578 2020-07-10 eess.IV cs.CV 57%

A Vision-based Social Distancing and Critical Density Detection System for COVID-19

Dongfang Yang, Ekim Yurtsever, Vishnu Renganathan, Keith A. Redmill, Ümit Özgüner

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.00809 2020-07-03 eess.AS cs.SD 57%

Automated Empathy Detection for Oncology Encounters

Zhuohao Chen, James Gibson, Ming-Chang Chiu, Qiaohong Hu, Tara K Knight, Daniella Meeker, James A Tulsky, Kathryn I Pollak, Shrikanth Narayanan

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by the 8TH IEEE International Conference on Healthcare Informatics (ICHI2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.12041 2020-06-23 cs.CV 57%

Characterizing Hirability via Personality and Behavior

Harshit Malik, Hersh Dhillon, Roland Goecke, Ramanathan Subramanian

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.01595 2020-06-03 eess.AS cs.LG cs.SD eess.IV stat.ML 57%

Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data

Haytham M. Fayek, Anurag Kumar

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments 29th International Joint Conference on Artificial Intelligence (IJCAI 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.04911 2020-04-13 cs.CV 57%

Analyze and Development System with Multiple Biometric Identification

Sher Dadakhanov

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Multiple Biometric Identification

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.00369 2020-04-08 cs.NI cs.MM 57%

Demonstrating Immersive Media Delivery on 5G Broadcast and Multicast Testing Networks

De Mi, Joe Eyles, Tero Jokela, Swen Petersen, Roman Odarchenko, Ece Ozturk, Duy-Kha Chau, Tuan Tran, Rory Turnbull, Heikki Kokkinen, Baruch Altman, Menno Bot, Darko Ratkaj, Olaf Renner, David Gomez-Barquero, Jordi Joan Gimenez

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

Comments 16 pages, 22 figures, IEEE Trans. Broadcasting

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.00825 2020-03-03 cs.CV eess.IV 57%

SIP-SegNet: A Deep Convolutional Encoder-Decoder Network for Joint Semantic Segmentation and Extraction of Sclera, Iris and Pupil based on Periocular Region Suppression

Bilal Hassan, Ramsha Ahmed, Taimur Hassan, Naoufel Werghi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.09414 2020-01-28 cs.CV 57%

Curriculum Audiovisual Learning

Di Hu, Zheng Wang, Haoyi Xiong, Dong Wang, Feiping Nie, Dejing Dou

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.09919 2020-01-24 cs.HC cs.CV cs.SD 57%

Speech, Head, and Eye-based Cues for Continuous Affect Prediction

Jonny O'Dwyer

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted paper (pre-print) for 2019 8th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.06206 2020-01-20 cs.CL 57%

Multi-step Joint-Modality Attention Network for Scene-Aware Dialogue System

Yun-Wei Chu, Kuan-Yen Lin, Chao-Chun Hsu, Lun-Wei Ku

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments DSTC8 collocated with AAAI2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.05080 2020-01-16 cs.CV 57%

Automated Anonymisation of Visual and Audio Data in Classroom Studies

Ömer Sümer, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments The Workshops of the Thirty-Fourth AAAI Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.05920 2019-12-13 eess.AS cs.LG cs.SD stat.ML 57%

Measuring Mother-Infant Emotions By Audio Sensing

Xuewen Yao, Dong He, Tiancheng Jing, Kaya de Barbaro

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.02615 2019-12-06 eess.AS cs.LG cs.SD stat.ML 57%

Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events

Wim Boes, Hugo Van hamme

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Journal ref Proceedings of the 27th ACM International Conference on Multimedia (MM '19). ACM, New York, NY, USA, 1961-1969

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.00087 2019-12-03 eess.SP eess.AS 57%

Effects of a Hovering Unmanned Aerial Vehicle on Urban Soundscapes Perception

Antonio J. Torija, Zhengguang Li, Rod H. Self

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.12851 2019-12-02 cs.AI cs.LG 57%

Playing Games in the Dark: An approach for cross-modality transfer in reinforcement learning

Rui Silva, Miguel Vasco, Francisco S. Melo, Ana Paiva, Manuela Veloso

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.09649 2019-11-22 cs.CV 57%

Learning to Localize Sound Sources in Visual Scenes: Analysis and Applications

Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, In So Kweon

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments To appear in TPAMI. arXiv admin note: substantial text overlap with arXiv:1803.03849

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.02001 2019-11-06 cs.CV 57%

Dancing to Music

Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, Jan Kautz

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2019; Project page: https://github.com/NVlabs/Dancing2Music

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.05204 2019-10-23 cs.SD cs.LG eess.AS 57%

Acoustic Scene Classification by Implicitly Identifying Distinct Sound Events

Hongwei Song, Jiqing Han, Shiwen Deng, Zhihao Du

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments code URL typo, code is available at https://github.com/hackerekcah/distinct-events-asc.git

Journal ref Proc. Interspeech 2019, 3860-3864

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.07254 2019-10-17 cs.LG cs.SD eess.AS stat.ML 57%

Audio-Conditioned U-Net for Position Estimation in Full Sheet Images

Florian Henkel, Rainer Kelz, Gerhard Widmer

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted at International Workshop on Reading Music Systems 2019 (WoRMS)

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.07147 2019-09-17 eess.IV eess.AS 57%

Alternative Visual Units for an Optimized Phoneme-Based Lipreading System

Helen Bear, Richard Harvey

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted and published in Applied Sciences, 22pgs plus appendices and references

Journal ref Applied. Sciences. 2019, 9(18), 3870

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.06749 2019-09-17 cs.RO cs.AI 57%

MuMMER: Socially Intelligent Human-Robot Interaction in Public Spaces

Mary Ellen Foster, Bart Craenen, Amol Deshmukh, Oliver Lemon, Emanuele Bastianelli, Christian Dondrup, Ioannis Papaioannou, Andrea Vanzo, Jean-Marc Odobez, Olivier Canévet, Yuanzhouhan Cao, Weipeng He, Angel Martínez-González, Petr Motlicek, Rémy Siegfried, Rachid Alami, Kathleen Belhassein, Guilhem Buisan, Aurélie Clodic, Amandine Mayima, Yoan Sallami, Guillaume Sarthou, Phani-Teja Singamaneni, Jules Waldhart, Alexandre Mazel, Maxime Caniot, Marketta Niemelä, Päivi Heikkilä, Hanna Lammi, Antti Tammela

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.05417 2019-09-13 cs.CV cs.LG eess.SP 57%

Deep User Identification Model with Multiple Biometrics

Hyoung-Kyu Song, Ebrahim AlAlkeem, Jaewoong Yun, Tae-Ho Kim, Tae-Ho Kim, Hyerin Yoo, Dasom Heo, Chan Yeob Yeun, Myungsu Chae

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted, CIKM 2019 Workshop on DTMBio

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.00692 2019-09-04 cs.CL 57%

Story-oriented Image Selection and Placement

Sreyasi Nag Chowdhury, Simon Razniewski, Gerhard Weikum

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.06147 2019-07-30 cs.MM eess.IV 57%

Grounding Object Detections With Transcriptions

Yasufumi Moriya, Ramon Sanabria, Florian Metze, Gareth J. F. Jones

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.09238 2019-07-23 cs.SD eess.AS 57%

Crowdsourcing a Dataset of Audio Captions

Samuel Lipping, Konstantinos Drossos, Tuomas Virtanen

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.01011 2019-07-03 cs.LG cs.CL stat.ML 57%

Learning Representations from Imperfect Time Series Data via Tensor Rank Regularization

Paul Pu Liang, Zhun Liu, Yao-Hung Hubert Tsai, Qibin Zhao, Ruslan Salakhutdinov, Louis-Philippe Morency

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏