arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4587 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4587 篇

2305.17350 2023-05-30 cs.CL 74%

How Good is Automatic Segmentation as a Multimodal Discourse Annotation Aid?

Corbyn Terpstra, Ibrahim Khebour, Mariah Bradford, Brett Wisniewski, Nikhil Krishnaswamy, Nathaniel Blanchard

专题命中 音频语音多模态 :multimodal(title);分类 cs.CL

Comments 7 pages, 1 figure, 2 tables, Proceedings of 19th Joint ISO-ACL Workshop on Interoperable Semantic Annotation (ISA 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.01320 2023-03-27 cs.CV 74%

Synthesizing Photorealistic Virtual Humans Through Cross-modal Disentanglement

Siddarth Ravichandran, Ondřej Texler, Dimitar Dinev, Hyun Jae Kang

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.07315 2023-02-16 eess.AS cs.LG cs.SD 74%

A dataset for Audio-Visual Sound Event Detection in Movies

Rajat Hebbar, Digbalay Bose, Krishna Somandepalli, Veena Vijai, Shrikanth Narayanan

专题命中 音频语音多模态 :audio-visual(title);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16206 2022-11-30 cs.CV 74%

Exploring adaptation of VideoMAE for Audio-Visual Diarization & Social @ Ego4d Looking at me Challenge

Yinan He, Guo Chen

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08878 2022-11-17 cs.MM 74%

Video-Music Retrieval:A Dual-Path Cross-Modal Network

Xin Gu, Yinghua Shen, Chaohui Lv

专题命中 音频语音多模态 :cross-modal(title);分类 cs.MM

Comments 5pages,3figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16644 2022-11-01 cs.CV 74%

Unsupervised Audio-Visual Lecture Segmentation

Darshan Singh S, Anchit Gupta, C. V. Jawahar, Makarand Tapaswi

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

Comments 17 pages, 14 figures, 14 tables, Accepted to WACV 2023. Project page: https://cvit.iiit.ac.in/research/projects/cvit-projects/avlectures

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.00344 2022-11-01 cs.CV cs.HC cs.LG 74%

Towards Intercultural Affect Recognition: Audio-Visual Affect Recognition in the Wild Across Six Cultures

Leena Mathur, Ralph Adolphs, Maja J Matarić

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

Comments Accepted at IEEE International Conference on Automatic Face and Gesture Recognition (FG 2023), publication and presentation at refereed IEEE workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02755 2022-10-07 cs.CV 74%

Audio-Visual Face Reenactment

Madhav Agarwal, Rudrabha Mukhopadhyay, Vinay Namboodiri, C V Jawahar

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

Comments Winter Conference on Applications of Computer Vision (WACV), 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.12895 2022-07-27 eess.AS cs.SD 74%

Multimodal Speech Emotion Recognition using Cross Attention with Aligned Audio and Text

Yoonhyung Lee, Seunghyun Yoon, Kyomin Jung

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Comments 5 pages, accepted by INTERSPEECH 2020

Journal ref Proc. Interspeech 2020, 2717-2721

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00657 2022-07-12 eess.AS cs.SD 74%

Multimodal Clustering with Role Induced Constraints for Speaker Diarization

Nikolaos Flemotomos, Shrikanth Narayanan

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Comments To appear at Interspeech 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.14850 2022-05-31 cs.RO cs.LG cs.SD eess.AS 74%

Play it by Ear: Learning Skills amidst Occlusion through Audio-Visual Imitation Learning

Maximilian Du, Olivia Y. Lee, Suraj Nair, Chelsea Finn

专题命中 音频语音多模态 :audio-visual(title);分类 eess.AS

Journal ref Robotics Science and Systems (RSS) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13161 2022-03-25 cs.CV 74%

Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation

Xian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu, Rui Qian, Xinyi Lin, Xiaowei Zhou, Wayne Wu, Bo Dai, Bolei Zhou

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CV

Comments Accepted by IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2022. Camera-Ready Version, 19 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.09750 2022-02-22 cs.SD cs.IR cs.LG eess.AS 74%

Enhancing Affective Representations of Music-Induced EEG through Multimodal Supervision and latent Domain Adaptation

Kleanthis Avramidis, Christos Garoufis, Athanasia Zlatintsi, Petros Maragos

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Comments 5 pages, 3 figures, IEEE ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.05762 2022-01-17 cs.HC cs.LG cs.MM 74%

Multimodal analysis of the predictability of hand-gesture properties

Taras Kucherenko, Rajmund Nagy, Michael Neff, Hedvig Kjellström, Gustav Eje Henter

专题命中 音频语音多模态 :multimodal(title);分类 cs.MM

Comments Accepted at the International Conference on Autonomous Agents and Multiagent Systems (AAMAS) 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.08970 2022-01-04 eess.AS cs.LG cs.SD 74%

Disentanglement Learning for Variational Autoencoders Applied to Audio-Visual Speech Enhancement

Guillaume Carbajal, Julius Richter, Timo Gerkmann

专题命中 音频语音多模态 :audio-visual(title);分类 eess.AS

Comments arXiv admin note: text overlap with arXiv:2102.06454

Journal ref 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.01300 2021-11-04 cs.CV 74%

Masking Modalities for Cross-modal Video Retrieval

Valentin Gabeur, Arsha Nagrani, Chen Sun, Karteek Alahari, Cordelia Schmid

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CV

Comments Accepted at WACV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.14509 2021-10-05 cs.CV 74%

Families In Wild Multimedia: A Multimodal Database for Recognizing Kinship

Joseph P. Robinson, Zaid Khan, Yu Yin, Ming Shao, Yun Fu

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV

Journal ref IEEE Transactions on Multimedia (2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01926 2021-09-07 cs.CV 74%

Audio-Visual Transformer Based Crowd Counting

Usman Sajid, Xiangyu Chen, Hasan Sajid, Taejoon Kim, Guanghui Wang

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06401 2021-08-17 cs.SD eess.AS 74%

Cross-modal Spectrum Transformation Network For Acoustic Scene classification

Yang Liu, Alexandros Neophytou, Sunando Sengupta, Eric Sommerlade

专题命中 音频语音多模态 :cross-modal(title);分类 eess.AS

Journal ref ICASSP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.10380 2021-06-21 cs.CL 74%

End-to-end Speech Translation via Cross-modal Progressive Training

Rong Ye, Mingxuan Wang, Lei Li

专题命中 音频语音多模态 :cross-modal(title);分类 cs.CL

Comments Accepted at InterSpeech2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.04309 2021-05-11 cs.SD cs.LG eess.AS 74%

Multi-modal Conditional Bounding Box Regression for Music Score Following

Florian Henkel, Gerhard Widmer

专题命中 音频语音多模态 :multi-modal(title);分类 eess.AS

Comments Accepted for publication in the Proceedings of the 29th European Signal Processing Conference (EUSIPCO), Dublin, Ireland, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.09805 2021-04-20 cs.LG cs.CV 74%

Active Contrastive Learning of Audio-Visual Video Representations

Shuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale Song

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2006.13662 2021-03-02 cs.CV cs.LG 74%

Labelling unlabelled videos from scratch with multi-modal self-supervision

Yuki M. Asano, Mandela Patrick, Christian Rupprecht, Andrea Vedaldi

专题命中 音频语音多模态 :multi-modal(title);分类 cs.CV

Comments Accepted to NeurIPS 2020. Project page: https://www.robots.ox.ac.uk/~vgg/research/selavi, code: https://github.com/facebookresearch/selavi

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.03692 2020-11-03 cs.LG cs.SD eess.AS 74%

An Audio-Video Deep and Transfer Learning Framework for Multimodal Emotion Recognition in the wild

Denis Dresvyanskiy, Elena Ryumina, Heysem Kaya, Maxim Markitantov, Alexey Karpov, Wolfgang Minker

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Comments Results on test dataset and acknowledgements were added

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.14171 2020-10-28 cs.SD cs.IR cs.LG eess.AS stat.ML 74%

Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags

Xavier Favory, Konstantinos Drossos, Tuomas Virtanen, Xavier Serra

专题命中 音频语音多模态 :cross-modal(title);分类 eess.AS

Comments 5 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.06711 2020-08-04 cs.CV cs.LG cs.SD 74%

Emotions Don't Lie: An Audio-Visual Deepfake Detection Method Using Affective Cues

Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, Dinesh Manocha

专题命中 音频语音多模态 :audio-visual(title);分类 cs.CV

Comments Accepted to ACMMM-2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.01440 2020-02-13 eess.AS cs.SD eess.SP 74%

Audio-Visual Calibration with Polynomial Regression for 2-D Projection Using SVD-PHAT

Francois Grondin, Hao Tang, James Glass

专题命中 音频语音多模态 :audio-visual(title);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.13799 2020-02-12 eess.AS cs.LG cs.SD 74%

Multimodal Learning For Classroom Activity Detection

Hang Li, Yu Kang, Wenbiao Ding, Song Yang, Songfan Yang, Gale Yan Huang, Zitao Liu

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Comments The 45th International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.11927 2019-11-28 eess.AS 74%

Automatic prediction of suicidal risk in military couples using multimodal interaction cues from couples conversations

Sandeep Nallan Chakravarthula, Md Nasir, Shao-Yen Tseng, Haoqi Li, Tae Jin Park, Brian Baucom, Craig J. Bryan, Shrikanth Narayanan, Panayiotis Georgiou

专题命中 音频语音多模态 :multimodal(title);分类 eess.AS

Comments submitted to ICASSP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.12622 2019-09-30 cs.CL cs.CY cs.HC 74%

Multi-Modal Citizen Science: From Disambiguation to Transcription of Classical Literature

Maryam Foradi, Jan Kaßel, Johannes Pein, Gregory R. Crane

专题命中 音频语音多模态 :multi-modal(title);分类 cs.CL

Journal ref Proceedings of the 30th ACM Conference on Hypertext and Social Media - HT 2019, 49-53. Hof, Germany: ACM Press

详情

展开后加载摘要…

URL PDF HTML 收藏