arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46352 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4587 篇

2306.17404 2023-07-03 cs.CV 74%

QuAVF: Quality-aware Audio-Visual Fusion for Ego4D Talking to Me Challenge

Hsi-Che Lin, Chien-Yi Wang, Min-Hung Chen, Szu-Wei Fu, Yu-Chiang Frank Wang

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

Comments 1st place at Ego4D Talking to Me (TTM) Challenge 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.04924 2023-06-26 cs.LG cs.AI 74%

Bayesian Networks for the robust and unbiased prediction of depression and its symptoms utilizing speech and multimodal data

Salvatore Fara, Orlaith Hickey, Alexandra Georgescu, Stefano Goria, Emilia Molimpakis, Nicholas Cummins

专题命中 音频语音多模态 :multimodal(title);分类 cs.AI

Comments Accepted for publication at Interspeech 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17350 2023-05-30 cs.CL 74%

How Good is Automatic Segmentation as a Multimodal Discourse Annotation Aid?

Corbyn Terpstra, Ibrahim Khebour, Mariah Bradford, Brett Wisniewski, Nikhil Krishnaswamy, Nathaniel Blanchard

专题命中 音频语音多模态 :multimodal(title);分类 cs.CL

Comments 7 pages, 1 figure, 2 tables, Proceedings of 19th Joint ISO-ACL Workshop on Interoperable Semantic Annotation (ISA 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.01320 2023-03-27 cs.CV 74%

Synthesizing Photorealistic Virtual Humans Through Cross-modal Disentanglement

Siddarth Ravichandran, Ondřej Texler, Dimitar Dinev, Hyun Jae Kang

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.07315 2023-02-16 eess.AS cs.LG cs.SD 74%

A dataset for Audio-Visual Sound Event Detection in Movies

Rajat Hebbar, Digbalay Bose, Krishna Somandepalli, Veena Vijai, Shrikanth Narayanan

专题命中 音频语音多模态 :audio-visual(title);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16206 2022-11-30 cs.CV 74%

Exploring adaptation of VideoMAE for Audio-Visual Diarization & Social @ Ego4d Looking at me Challenge

Yinan He, Guo Chen

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08878 2022-11-17 cs.MM 74%

Video-Music Retrieval:A Dual-Path Cross-Modal Network

Xin Gu, Yinghua Shen, Chaohui Lv

专题命中 音频语音多模态 :cross-modal(title);分类 cs.MM

Comments 5pages,3figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16644 2022-11-01 cs.CV 74%

Unsupervised Audio-Visual Lecture Segmentation

Darshan Singh S, Anchit Gupta, C. V. Jawahar, Makarand Tapaswi

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

Comments 17 pages, 14 figures, 14 tables, Accepted to WACV 2023. Project page: https://cvit.iiit.ac.in/research/projects/cvit-projects/avlectures

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.00344 2022-11-01 cs.CV cs.HC cs.LG 74%

Towards Intercultural Affect Recognition: Audio-Visual Affect Recognition in the Wild Across Six Cultures

Leena Mathur, Ralph Adolphs, Maja J Matarić

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

Comments Accepted at IEEE International Conference on Automatic Face and Gesture Recognition (FG 2023), publication and presentation at refereed IEEE workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02755 2022-10-07 cs.CV 74%

Audio-Visual Face Reenactment

Madhav Agarwal, Rudrabha Mukhopadhyay, Vinay Namboodiri, C V Jawahar

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

Comments Winter Conference on Applications of Computer Vision (WACV), 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12895 2022-07-27 eess.AS cs.SD 74%

Multimodal Speech Emotion Recognition using Cross Attention with Aligned Audio and Text

Yoonhyung Lee, Seunghyun Yoon, Kyomin Jung

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Comments 5 pages, accepted by INTERSPEECH 2020

Journal ref Proc. Interspeech 2020, 2717-2721

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00657 2022-07-12 eess.AS cs.SD 74%

Multimodal Clustering with Role Induced Constraints for Speaker Diarization

Nikolaos Flemotomos, Shrikanth Narayanan

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Comments To appear at Interspeech 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.14850 2022-05-31 cs.RO cs.LG cs.SD eess.AS 74%

Play it by Ear: Learning Skills amidst Occlusion through Audio-Visual Imitation Learning

Maximilian Du, Olivia Y. Lee, Suraj Nair, Chelsea Finn

专题命中 音频语音多模态 :audio-visual(title);分类 eess.AS

Journal ref Robotics Science and Systems (RSS) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13161 2022-03-25 cs.CV 74%

Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation

Xian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu, Rui Qian, Xinyi Lin, Xiaowei Zhou, Wayne Wu, Bo Dai, Bolei Zhou

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CV

Comments Accepted by IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2022. Camera-Ready Version, 19 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.09750 2022-02-22 cs.SD cs.IR cs.LG eess.AS 74%

Enhancing Affective Representations of Music-Induced EEG through Multimodal Supervision and latent Domain Adaptation

Kleanthis Avramidis, Christos Garoufis, Athanasia Zlatintsi, Petros Maragos

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Comments 5 pages, 3 figures, IEEE ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.05762 2022-01-17 cs.HC cs.LG cs.MM 74%

Multimodal analysis of the predictability of hand-gesture properties

Taras Kucherenko, Rajmund Nagy, Michael Neff, Hedvig Kjellström, Gustav Eje Henter

专题命中 音频语音多模态 :multimodal(title);分类 cs.MM

Comments Accepted at the International Conference on Autonomous Agents and Multiagent Systems (AAMAS) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08970 2022-01-04 eess.AS cs.LG cs.SD 74%

Disentanglement Learning for Variational Autoencoders Applied to Audio-Visual Speech Enhancement

Guillaume Carbajal, Julius Richter, Timo Gerkmann

专题命中 音频语音多模态 :audio-visual(title);分类 eess.AS

Comments arXiv admin note: text overlap with arXiv:2102.06454

Journal ref 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.01300 2021-11-04 cs.CV 74%

Masking Modalities for Cross-modal Video Retrieval

Valentin Gabeur, Arsha Nagrani, Chen Sun, Karteek Alahari, Cordelia Schmid

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CV

Comments Accepted at WACV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.14509 2021-10-05 cs.CV 74%

Families In Wild Multimedia: A Multimodal Database for Recognizing Kinship

Joseph P. Robinson, Zaid Khan, Yu Yin, Ming Shao, Yun Fu

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV

Journal ref IEEE Transactions on Multimedia (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01926 2021-09-07 cs.CV 74%

Audio-Visual Transformer Based Crowd Counting

Usman Sajid, Xiangyu Chen, Hasan Sajid, Taejoon Kim, Guanghui Wang

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06401 2021-08-17 cs.SD eess.AS 74%

Cross-modal Spectrum Transformation Network For Acoustic Scene classification

Yang Liu, Alexandros Neophytou, Sunando Sengupta, Eric Sommerlade

专题命中 音频语音多模态 :cross-modal(title);分类 eess.AS

Journal ref ICASSP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10380 2021-06-21 cs.CL 74%

End-to-end Speech Translation via Cross-modal Progressive Training

Rong Ye, Mingxuan Wang, Lei Li

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CL

Comments Accepted at InterSpeech2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.04309 2021-05-11 cs.SD cs.LG eess.AS 74%

Multi-modal Conditional Bounding Box Regression for Music Score Following

Florian Henkel, Gerhard Widmer

专题命中 音频语音多模态 :multi-modal(title);分类 eess.AS

Comments Accepted for publication in the Proceedings of the 29th European Signal Processing Conference (EUSIPCO), Dublin, Ireland, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.09805 2021-04-20 cs.LG cs.CV 74%

Active Contrastive Learning of Audio-Visual Video Representations

Shuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale Song

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.13662 2021-03-02 cs.CV cs.LG 74%

Labelling unlabelled videos from scratch with multi-modal self-supervision

Yuki M. Asano, Mandela Patrick, Christian Rupprecht, Andrea Vedaldi

专题命中 音频语音多模态 :multi-modal(title);分类 cs.CV

Comments Accepted to NeurIPS 2020. Project page: https://www.robots.ox.ac.uk/~vgg/research/selavi, code: https://github.com/facebookresearch/selavi

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.03692 2020-11-03 cs.LG cs.SD eess.AS 74%

An Audio-Video Deep and Transfer Learning Framework for Multimodal Emotion Recognition in the wild

Denis Dresvyanskiy, Elena Ryumina, Heysem Kaya, Maxim Markitantov, Alexey Karpov, Wolfgang Minker

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Comments Results on test dataset and acknowledgements were added

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.14171 2020-10-28 cs.SD cs.IR cs.LG eess.AS stat.ML 74%

Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags

Xavier Favory, Konstantinos Drossos, Tuomas Virtanen, Xavier Serra

专题命中 音频语音多模态 :cross-modal(title);分类 eess.AS

Comments 5 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.06711 2020-08-04 cs.CV cs.LG cs.SD 74%

Emotions Don't Lie: An Audio-Visual Deepfake Detection Method Using Affective Cues

Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, Dinesh Manocha

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

Comments Accepted to ACMMM-2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.01440 2020-02-13 eess.AS cs.SD eess.SP 74%

Audio-Visual Calibration with Polynomial Regression for 2-D Projection Using SVD-PHAT

Francois Grondin, Hao Tang, James Glass

专题命中 音频语音多模态 :audio-visual(title);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.13799 2020-02-12 eess.AS cs.LG cs.SD 74%

Multimodal Learning For Classroom Activity Detection

Hang Li, Yu Kang, Wenbiao Ding, Song Yang, Songfan Yang, Gale Yan Huang, Zitao Liu

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Comments The 45th International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏