arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2010.02015 2020-10-06 cs.MM cs.GR 57%

Combined Hapto-Visual and Auditory Rendering of Cultural Heritage Objects

Praseedha Krishnan Aniyath, Sreeni Kamalalayam Gopalan, Priyadarshini K, Subhasis Chaudhuri

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.MM

Comments Accepted to ACCVw 2014

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.07381 2020-09-22 cs.RO cs.AI 57%

Spatial Concept-Based Navigation with Human Speech Instructions via Probabilistic Inference on Bayesian Generative Model

Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi, Tetsunari Inamura

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted to Advanced Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.02119 2020-09-07 cs.GR cs.CV cs.HC 57%

Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity

Youngwoo Yoon, Bok Cha, Joo-Haeng Lee, Minsu Jang, Jaeyeon Lee, Jaehong Kim, Geehyuk Lee

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 16 pages; ACM Transactions on Graphics (SIGGRAPH Asia 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.12855 2020-09-01 cs.MM cs.CY cs.HC 57%

Personal Food Model

Ali Rostami, Vaibhav Pandey, Nitish Nag, Vesper Wang, Ramesh Jain

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM

Journal ref Proceedings of the 28th ACM International Conference on Multimedia (MM '20), October 12--16, 2020, Seattle, WA, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.05023 2020-08-13 cs.CV 57%

Audio- and Gaze-driven Facial Animation of Codec Avatars

Alexander Richard, Colin Lea, Shugao Ma, Juergen Gall, Fernando de la Torre, Yaser Sheikh

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.04617 2020-08-12 eess.AS cs.SD 57%

Alzheimer's Dementia Detection from Audio and Text Modalities

Edward L. Campbell, Laura Docío-Fernández, Javier Jiménez Raboso, Carmen García-Mateo

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 5 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.03781 2020-08-11 cs.CV 57%

SemEval-2020 Task 8: Memotion Analysis -- The Visuo-Lingual Metaphor!

Chhavi Sharma, Deepesh Bhageria, William Scott, Srinivas PYKL, Amitava Das, Tanmoy Chakraborty, Viswanath Pulabaigari, Bjorn Gamback

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.02070 2020-08-06 eess.AS cs.LG cs.SD 57%

Content based singing voice source separation via strong conditioning using aligned phonemes

Gabriel Meseguer-Brocal, Geoffroy Peeters

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 21st International Society for Music Information Retrieval Conference 11-15 October 2020, Montreal, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.03889 2020-08-03 eess.AS cs.SD 57%

Neural Spatio-Temporal Beamformer for Target Speech Separation

Yong Xu, Meng Yu, Shi-Xiong Zhang, Lianwu Chen, Chao Weng, Jianming Liu, Dong Yu

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments accepted to Interspeech2020, Demo: https://yongxuustc.github.io/mtmvdr/

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.08874 2020-07-24 eess.AS cs.LG cs.SD 57%

Speech Emotion Recognition with Dual-Sequence LSTM Architecture

Jianyou Wang, Michael Xue, Ryan Culhane, Enmao Diao, Jie Ding, Vahid Tarokh

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by ICASSP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.06355 2020-07-15 cs.CV 57%

Multiple Sound Sources Localization from Coarse to Fine

Rui Qian, Di Hu, Heinrich Dinkel, Mengyue Wu, Ning Xu, Weiyao Lin

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments to appear in ECCV 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.03578 2020-07-10 eess.IV cs.CV 57%

A Vision-based Social Distancing and Critical Density Detection System for COVID-19

Dongfang Yang, Ekim Yurtsever, Vishnu Renganathan, Keith A. Redmill, Ümit Özgüner

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.00809 2020-07-03 eess.AS cs.SD 57%

Automated Empathy Detection for Oncology Encounters

Zhuohao Chen, James Gibson, Ming-Chang Chiu, Qiaohong Hu, Tara K Knight, Daniella Meeker, James A Tulsky, Kathryn I Pollak, Shrikanth Narayanan

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by the 8TH IEEE International Conference on Healthcare Informatics (ICHI2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.12041 2020-06-23 cs.CV 57%

Characterizing Hirability via Personality and Behavior

Harshit Malik, Hersh Dhillon, Roland Goecke, Ramanathan Subramanian

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.01595 2020-06-03 eess.AS cs.LG cs.SD eess.IV stat.ML 57%

Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data

Haytham M. Fayek, Anurag Kumar

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments 29th International Joint Conference on Artificial Intelligence (IJCAI 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.04911 2020-04-13 cs.CV 57%

Analyze and Development System with Multiple Biometric Identification

Sher Dadakhanov

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Multiple Biometric Identification

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.00369 2020-04-08 cs.NI cs.MM 57%

Demonstrating Immersive Media Delivery on 5G Broadcast and Multicast Testing Networks

De Mi, Joe Eyles, Tero Jokela, Swen Petersen, Roman Odarchenko, Ece Ozturk, Duy-Kha Chau, Tuan Tran, Rory Turnbull, Heikki Kokkinen, Baruch Altman, Menno Bot, Darko Ratkaj, Olaf Renner, David Gomez-Barquero, Jordi Joan Gimenez

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

Comments 16 pages, 22 figures, IEEE Trans. Broadcasting

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.00825 2020-03-03 cs.CV eess.IV 57%

SIP-SegNet: A Deep Convolutional Encoder-Decoder Network for Joint Semantic Segmentation and Extraction of Sclera, Iris and Pupil based on Periocular Region Suppression

Bilal Hassan, Ramsha Ahmed, Taimur Hassan, Naoufel Werghi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.09414 2020-01-28 cs.CV 57%

Curriculum Audiovisual Learning

Di Hu, Zheng Wang, Haoyi Xiong, Dong Wang, Feiping Nie, Dejing Dou

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.09919 2020-01-24 cs.HC cs.CV cs.SD 57%

Speech, Head, and Eye-based Cues for Continuous Affect Prediction

Jonny O'Dwyer

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted paper (pre-print) for 2019 8th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.06206 2020-01-20 cs.CL 57%

Multi-step Joint-Modality Attention Network for Scene-Aware Dialogue System

Yun-Wei Chu, Kuan-Yen Lin, Chao-Chun Hsu, Lun-Wei Ku

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments DSTC8 collocated with AAAI2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.05080 2020-01-16 cs.CV 57%

Automated Anonymisation of Visual and Audio Data in Classroom Studies

Ömer Sümer, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments The Workshops of the Thirty-Fourth AAAI Conference on Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.05920 2019-12-13 eess.AS cs.LG cs.SD stat.ML 57%

Measuring Mother-Infant Emotions By Audio Sensing

Xuewen Yao, Dong He, Tiancheng Jing, Kaya de Barbaro

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.02615 2019-12-06 eess.AS cs.LG cs.SD stat.ML 57%

Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events

Wim Boes, Hugo Van hamme

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Journal ref Proceedings of the 27th ACM International Conference on Multimedia (MM '19). ACM, New York, NY, USA, 1961-1969

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.00087 2019-12-03 eess.SP eess.AS 57%

Effects of a Hovering Unmanned Aerial Vehicle on Urban Soundscapes Perception

Antonio J. Torija, Zhengguang Li, Rod H. Self

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.12851 2019-12-02 cs.AI cs.LG 57%

Playing Games in the Dark: An approach for cross-modality transfer in reinforcement learning

Rui Silva, Miguel Vasco, Francisco S. Melo, Ana Paiva, Manuela Veloso

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.09649 2019-11-22 cs.CV 57%

Learning to Localize Sound Sources in Visual Scenes: Analysis and Applications

Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, In So Kweon

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments To appear in TPAMI. arXiv admin note: substantial text overlap with arXiv:1803.03849

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.02001 2019-11-06 cs.CV 57%

Dancing to Music

Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, Jan Kautz

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments NeurIPS 2019; Project page: https://github.com/NVlabs/Dancing2Music

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.05204 2019-10-23 cs.SD cs.LG eess.AS 57%

Acoustic Scene Classification by Implicitly Identifying Distinct Sound Events

Hongwei Song, Jiqing Han, Shiwen Deng, Zhihao Du

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments code URL typo, code is available at https://github.com/hackerekcah/distinct-events-asc.git

Journal ref Proc. Interspeech 2019, 3860-3864

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.07254 2019-10-17 cs.LG cs.SD eess.AS stat.ML 57%

Audio-Conditioned U-Net for Position Estimation in Full Sheet Images

Florian Henkel, Rainer Kelz, Gerhard Widmer

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted at International Workshop on Reading Music Systems 2019 (WoRMS)

详情

展开后加载摘要…

URL PDF HTML 收藏