arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2205.04274 2022-05-31 cs.CL cs.AI cs.CV 67%

Detecting and Understanding Harmful Memes: A Survey

Shivam Sharma, Firoj Alam, Md. Shad Akhtar, Dimitar Dimitrov, Giovanni Da San Martino, Hamed Firooz, Alon Halevy, Fabrizio Silvestri, Preslav Nakov, Tanmoy Chakraborty

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at IJCAI-ECAI 2022 (Survey Track) - Editorial Feedback Revised, 9 pages (7 main + 2 reference pages)

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.02639 2022-05-16 cs.CV cs.CL cs.LG cs.SD eess.AS 67%

MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound

Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, Yejin Choi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.CL、eess.AS

Comments CVPR 2022. Project page at https://rowanzellers.com/merlotreserve

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08995 2022-05-04 cs.SD cs.CL cs.CV eess.AS 67%

Connecting the Dots between Audio and Text without Parallel Data through Visual Knowledge Transfer

Yanpeng Zhao, Jack Hessel, Youngjae Yu, Ximing Lu, Rowan Zellers, Yejin Choi

专题命中 音频语音多模态 :image-text(abstract);分类 cs.CV、cs.CL、eess.AS

Comments Accepted to NAACL 2022. Our code is available at https://github.com/zhaoyanpeng/vipant

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.14272 2022-05-02 cs.CL cs.AI cs.SD eess.AS 67%

End-to-end Spoken Conversational Question Answering: Task, Dataset and Model

Chenyu You, Nuo Chen, Fenglin Liu, Shen Ge, Xian Wu, Yuexian Zou

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments In Findings of NAACL 2022. arXiv admin note: substantial text overlap with arXiv:2010.08923

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09919 2022-04-22 cs.HC cs.AI cs.MM cs.SD eess.AS 67%

Sonic Interactions in Virtual Environments: the Egocentric Audio Perspective of the Digital Twin

Michele Geronazzo, Stefania Serafin

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、cs.MM、eess.AS

Comments 46 pages, 5 figures. Pre-print version of the introduction to the book "Sonic Interactions in Virtual Environments" in press for Springer's Human-Computer Interaction Series, Open Access license. The pre-print editors' copy of the book can be found at https://vbn.aau.dk/en/publications/sonic-interactions-in-virtual-environments - full book info: https://sive.create.aau.dk/index.php/sivebook/

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.01726 2022-04-06 cs.CV cs.AI eess.AS 67%

Lip to Speech Synthesis with Visual Context Attentional GAN

Minsu Kim, Joanna Hong, Yong Man Ro

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.AI、eess.AS

Comments Published at NeurIPS 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15991 2022-03-31 cs.CV cs.MM cs.SD eess.AS 67%

The Sound of Bounding-Boxes

Takashi Oya, Shohei Iwase, Shigeo Morishima

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

Comments 6 pages, 5 figures, ICPR (International Conference on Pattern Recognition) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.09324 2022-03-30 cs.CV cs.LG cs.MM cs.SD eess.AS 67%

Localizing Visual Sounds the Easy Way

Shentong Mo, Pedro Morgado

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.08243 2022-03-16 eess.AS cs.CL cs.CV cs.LG cs.SD eess.IV 67%

Neural Dubber: Dubbing for Videos According to Scripts

Chenxu Hu, Qiao Tian, Tingle Li, Yuping Wang, Yuxuan Wang, Hang Zhao

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、eess.AS

Comments Accepted by NeurIPS 2021; Project page at https://tsinghua-mars-lab.github.io/NeuralDubber/

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.10453 2022-02-23 cs.CV cs.LG cs.MM cs.SD eess.AS 67%

Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses

Phoebe Chua, Dimos Makris, Dorien Herremans, Gemma Roig, Kat Agres

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments 16 pages with 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.05846 2021-11-11 cs.SD cs.CV cs.MM cs.RO eess.AS 67%

Structure from Silence: Learning Scene Structure from Ambient Sound

Ziyang Chen, Xixi Hu, Andrew Owens

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted to CoRL 2021 (Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08052 2021-10-26 cs.CV cs.MM cs.RO cs.SD eess.AS 67%

The Boombox: Visual Reconstruction from Acoustic Vibrations

Boyuan Chen, Mia Chiquier, Hod Lipson, Carl Vondrick

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments CoRL 2021. Website: boombox.cs.columbia.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.09105 2021-09-22 cs.CL cs.AI cs.LG eess.AS 67%

What BERT Based Language Models Learn in Spoken Transcripts: An Empirical Study

Ayush Kumar, Mukuntha Narayanan Sundararaman, Jithendra Vepa

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Comments BlackboxNLP @ EMNLP 2021 (15 pages, includes Appendix)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.02607 2021-08-06 cs.CV cs.MM cs.SD eess.AS eess.IV 67%

UniCon: Unified Context Network for Robust Active Speaker Detection

Yuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu, Zhongqin Wu, Shiguang Shan, Xilin Chen

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

Comments 10 pages, 6 figures; to appear at ACM Multimedia 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.09262 2021-07-21 cs.LG cs.AI cs.CV cs.MM cs.SD 67%

FoleyGAN: Visually Guided Generative Adversarial Network-Based Synchronous Sound Generation in Silent Videos

Sanchita Ghose, John J. Prevost

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM

Comments This article is under review in IEEE Transaction on Multimedia. It contains total 12 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.02687 2021-04-07 cs.CV cs.AI cs.MM 67%

Strumming to the Beat: Audio-Conditioned Contrastive Video Textures

Medhini Narasimhan, Shiry Ginosar, Andrew Owens, Alexei A. Efros, Trevor Darrell

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM

Comments Project website at https://medhini.github.io/audio_video_textures/

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.02026 2021-04-06 cs.CV cs.MM cs.SD eess.AS 67%

Cyclic Co-Learning of Sounding Object Visual Grounding and Sound Separation

Yapeng Tian, Di Hu, Chenliang Xu

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

Comments CVPR 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.08606 2021-03-22 cs.CV cs.MM cs.SD eess.AS 67%

End-to-End Lip Synchronisation Based on Pattern Classification

You Jin Kim, Hee Soo Heo, Soo-Whan Chung, Bong-Jin Lee

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments slt 2021 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.01400 2021-03-19 eess.AS cs.CL cs.CV cs.LG 67%

Does Visual Self-Supervision Improve Learning of Speech Representations for Emotion Recognition?

Abhinav Shukla, Stavros Petridis, Maja Pantic

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.CL、eess.AS

Comments Accepted for publication in IEEE Transactions on Affective Computing; v3: Publication-ready version including additional experiments and discussion

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.15635 2021-01-01 cs.MM cs.AI cs.CV 67%

Leveraging Audio Gestalt to Predict Media Memorability

Lorin Sweeney, Graham Healy, Alan F. Smeaton

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 3 pages, 1 Figure, 2 Tables

Journal ref MediaEval Multimedia Benchmark Workshop Working Notes, 14-15 December 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.08474 2020-12-22 eess.SP cs.CY cs.HC stat.AP 67%

TILES-2018, a longitudinal physiologic and behavioral data set of hospital workers

Karel Mundnich, Brandon M. Booth, Michelle L'Hommedieu, Tiantian Feng, Benjamin Girault, Justin L'Hommedieu, Mackenzie Wildman, Sophia Skaaden, Amrutha Nadarajan, Jennifer L. Villatte, Tiago H. Falk, Kristina Lerman, Emilio Ferrara, Shrikanth Narayanan

专题命中 音频语音多模态 :multimodal(abstract);multi-modal(abstract)

Comments 57 pages, 9 figures, journal paper

Journal ref Sci Data 7, 354 (2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.05103 2020-09-14 cs.CV cs.MM cs.SD eess.AS 67%

Emotion-Based End-to-End Matching Between Image and Music in Valence-Arousal Space

Sicheng Zhao, Yaxian Li, Xingxu Yao, Weizhi Nie, Pengfei Xu, Jufeng Yang, Kurt Keutzer

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted by ACM Multimedia 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.09902 2020-07-21 cs.CV cs.MM cs.SD eess.AS 67%

Sep-Stereo: Visually Guided Stereophonic Audio Generation by Associating Source Separation

Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, Ziwei Liu

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

Comments To appear in Proceedings of the European Conference on Computer Vision (ECCV), 2020. Code, models, and video results are available on our webpage: https://hangz-nju-cuhk.github.io/projects/Sep-Stereo

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.09476 2020-04-21 cs.CV cs.LG cs.MM cs.SD eess.AS 67%

Music Gesture for Visual Sound Separation

Chuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum, Antonio Torralba

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.MM、eess.AS

Comments CVPR 2020. Project page: http://music-gesture.csail.mit.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.00418 2020-03-03 cs.CV cs.AI cs.LG cs.MM cs.SD 67%

Towards Automatic Face-to-Face Translation

Prajwal K R, Rudrabha Mukhopadhyay, Jerin Philip, Abhishek Jha, Vinay Namboodiri, C. V. Jawahar

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments 9 pages (including references), 5 figures, Published in ACM Multimedia, 2019

Journal ref MM '19: Proceedings of the 27th ACM International Conference on Multimedia; October 2019; Pages 1428-1436

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.05639 2020-02-19 cs.CL cs.MM eess.AS 67%

Looking Enhances Listening: Recovering Missing Speech Using Images

Tejas Srinivasan, Ramon Sanabria, Florian Metze

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.MM、eess.AS

Comments Accepted to ICASSP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.01990 2019-10-07 cs.CL cs.AI eess.AS 67%

Detecting Deception in Political Debates Using Acoustic and Textual Features

Daniel Kopev, Ahmed Ali, Ivan Koychev, Preslav Nakov

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

Journal ref ASRU-2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1907.01154 2019-07-03 cs.MM cs.AI cs.SD eess.AS 67%

Adaptive Music Composition for Games

Patrick Hutchings, Jon McCormack

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、cs.MM、eess.AS

Comments Preprint. Accepted for publication in IEEE Transactions on Games, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1809.02587 2018-09-10 cs.SD cs.CV cs.LG cs.MM eess.AS 67%

Self-Supervised Generation of Spatial Audio for 360 Video

Pedro Morgado, Nuno Vasconcelos, Timothy Langlois, Oliver Wang

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、cs.MM、eess.AS

Comments To appear in NIPS 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.06652 2018-05-18 cs.HC 67%

Affective computing using speech and eye gaze: a review and bimodal system proposal for continuous affect prediction

Jonny O'Dwyer, Niall Murray, Ronan Flynn

专题命中 音频语音多模态 :multi-modal(abstract);audio-visual(abstract)

Comments Submitted to International Journal of Human-Computer Studies

详情

展开后加载摘要…

URL PDF HTML 收藏