arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4585 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4585 篇

2301.13190 2023-01-31 cs.CV 79%

Audio-Visual Segmentation with Semantics

Jinxing Zhou, Xuyang Shen, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, Yiran Zhong

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Submitted to TPAMI as a journal extension of ECCV 2022. Jinxing Zhou, Xuyang Shen, and Jianyuan Wang contribute equally to this work. Meng Wang and Yiran Zhong are the corresponding authors. Code is available at https://github.com/OpenNLPLab/AVSBench. Online benchmark is available at http://www.avlbench.opennlplab.cn. arXiv admin note: substantial text overlap with arXiv:2207.05042

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00145 2023-01-03 cs.CV 79%

Attentional Graph Convolutional Network for Structure-aware Audio-Visual Scene Classification

Liguang Zhou, Yuhongze Zhou, Xiaonan Qi, Junjie Hu, Tin Lun Lam, Yangsheng Xu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.06887 2022-11-28 cs.HC cs.CV cs.LG 79%

AVCAffe: A Large Scale Audio-Visual Dataset of Cognitive Load and Affect for Remote Work

Pritam Sarkar, Aaron Posen, Ali Etemad

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted in AAAI 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.00970 2022-11-23 eess.AS cs.SD 79%

Self-supervised Learning of Audio Representations from Audio-Visual Data using Spatial Alignment

Shanshan Wang, Archontis Politis, Annamaria Mesaros, Tuomas Virtanen

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.10885 2022-11-22 cs.SD eess.AS 79%

Contrastive Regularization for Multimodal Emotion Recognition Using Audio and Text

Fan Qian, Jiqing Han

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Completed in October 2020 and submitted to ICASSP2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.09980 2022-11-21 cs.CV 79%

Contrastive Positive Sample Propagation along the Audio-Visual Event Line

Jinxing Zhou, Dan Guo, Meng Wang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted to TPAMI; Dataset and Code are available at https://github.com/jasongief/CPSP. arXiv admin note: substantial text overlap with arXiv:2104.00239

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.05543 2022-11-11 cs.SD cs.LG eess.AS 79%

Vis2Mus: Exploring Multimodal Representation Mapping for Controllable Music Generation

Runbang Zhang, Yixiao Zhang, Kai Shao, Ying Shan, Gus Xia

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Submitted to ICASSP 2023. GitHub repo: https://github.com/ldzhangyx/vis2mus

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.15385 2022-10-28 eess.AS cs.SD eess.SP 79%

Self-Supervised Training of Speaker Encoder with Multi-Modal Diverse Positive Pairs

Ruijie Tao, Kong Aik Lee, Rohan Kumar Das, Ville Hautamäki, Haizhou Li

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.03051 2022-10-25 cs.RO cs.AI cs.LG 79%

Perceive, Represent, Generate: Translating Multimodal Information to Robotic Motion Trajectories

Fábio Vital, Miguel Vasco, Alberto Sardinha, Francisco Melo

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments 14 pages, 4 figures, 8 tables, 1 algorithm

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.07846 2022-10-18 cs.LG cs.SD eess.AS q-bio.BM 79%

Telehealthcare and Telepathology in Pandemic: A Noninvasive, Low-Cost Micro-Invasive and Multimodal Real-Time Online Application for Early Diagnosis of COVID-19 Infection

Abdullah Bin Shams, Md. Mohsin Sarker Raihan, Md. Mohi Uddin Khan, Ocean Monjur, Rahat Bin Preo

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments 32 Pages. This article has been submitted for review to a prestigious journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.04063 2022-10-12 eess.AS cs.SD 79%

LiMuSE: Lightweight Multi-modal Speaker Extraction

Qinghua Liu, Yating Huang, Yunzhe Hao, Jiaming Xu, Bo Xu

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments Accepted to IEEE SLT 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.06268 2022-09-28 eess.AS cs.SD 79%

AMFFCN: Attentional Multi-layer Feature Fusion Convolution Network for Audio-visual Speech Enhancement

Xinmeng Xu, Jianjun Hao

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments arXiv admin note: text overlap with arXiv:2101.05975

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.00257 2022-09-21 cs.CL 79%

Sentiment Word Aware Multimodal Refinement for Multimodal Sentiment Analysis with ASR Errors

Yang Wu, Yanyan Zhao, Hao Yang, Song Chen, Bing Qin, Xiaohuan Cao, Wenting Zhao

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Findings of ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.09890 2022-09-19 eess.AS cs.LG cs.SD 79%

Multi-Modal Pre-Training for Automated Speech Recognition

David M. Chan, Shalini Ghosh, Debmalya Chakrabarty, Björn Hoffmeister

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments Presented at ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.03137 2022-09-08 cs.LG cs.AI 79%

Federated Transfer Learning with Multimodal Data

Yulian Sun

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments 73 pages, 54 figures, master thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08909 2022-08-24 cs.HC cs.CL cs.CY 79%

"Are you okay, honey?": Recognizing Emotions among Couples Managing Diabetes in Daily Life using Multimodal Real-World Smartwatch Data

George Boateng, Xiangyu Zhao, Malgorzata Speichert, Elgar Fleisch, Janina Lüscher, Theresa Pauly, Urte Scholz, Guy Bodenmann, Tobias Kowatsch

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 19 pages, Under review at ACM IMWUT

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.08118 2022-08-18 cs.CV 79%

Extreme-scale Talking-Face Video Upsampling with Audio-Visual Priors

Sindhu B Hegde, Rudrabha Mukhopadhyay, Vinay P Namboodiri, C. V. Jawahar

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted in ACM-MM 2022, 10 pages, 6 pages supplementary, 18 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.01917 2022-08-04 cs.SD cs.HC cs.LG eess.AS 79%

Zero-Shot Style Transfer for Gesture Animation driven by Text and Speech using Adversarial Disentanglement of Multimodal Style Encoding

Mireille Fares, Michele Grimaldi, Catherine Pelachaud, Nicolas Obin

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.11573 2022-08-02 cs.CV 79%

Joint-Modal Label Denoising for Weakly-Supervised Audio-Visual Video Parsing

Haoyue Cheng, Zhaoyang Liu, Hang Zhou, Chen Qian, Wayne Wu, Limin Wang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by ECCV 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.06401 2022-07-12 cs.CV 79%

Self-supervised object detection from audio-visual correspondence

Triantafyllos Afouras, Yuki M. Asano, Francois Fagan, Andrea Vedaldi, Florian Metze

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted to CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.01832 2022-07-06 cs.SD eess.AS 79%

Glow-WaveGAN 2: High-quality Zero-shot Text-to-speech Synthesis and Any-to-any Voice Conversion

Yi Lei, Shan Yang, Jian Cong, Lei Xie, Dan Su

专题命中 音频语音多模态 :any-to-any(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.12771 2022-06-14 cs.CV 79%

Self-Supervised Moving Vehicle Detection from Audio-Visual Cues

Jannik Zürn, Wolfram Burgard

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.03885 2022-06-09 eess.AS eess.SP 79%

On the Integration of Acoustics and LiDAR: a Multi-Modal Approach to Acoustic Reflector Estimation

Ellen Riemens, Pablo Martínez-Nuevo, Jorge Martinez, Martin Møller, Richard C. Hendriks

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments 5 pages, 9 figures, to be published in EUSIPCO 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.14960 2022-05-24 cs.LG cs.AI 79%

Domino: Discovering Systematic Errors with Cross-Modal Embeddings

Sabri Eyuboglu, Maya Varma, Khaled Saab, Jean-Benoit Delbrouck, Christopher Lee-Messer, Jared Dunnmon, James Zou, Christopher Ré

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.AI

Comments ICLR 2022 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.05975 2022-05-24 eess.AS cs.SD eess.IV 79%

Multi-layer Feature Fusion Convolution Network for Audio-visual Speech Enhancement

Xinmeng Xu, Jianjun Hao

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.09744 2022-05-20 cs.LG cs.CY cs.MM 79%

Overcoming Language Disparity in Online Content Classification with Multimodal Learning

Gaurav Verma, Rohit Mujumdar, Zijie J. Wang, Munmun De Choudhury, Srijan Kumar

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments Accepted for publication at ICWSM 2022 as a full paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2008.09743 2022-05-16 cs.SD cs.CV 79%

A Efficient Multimodal Framework for Large Scale Emotion Recognition by Fusing Music and Electrodermal Activity Signals

Guanghao Yin, Shouqian Sun, Dian Yu, Dejian Li, Kejun Zhang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments ACM Transactions on Multimedia Computing, Communications, and Applications (Acceptance 07-Oct-2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00073 2022-05-03 cs.CV 79%

On Negative Sampling for Audio-Visual Contrastive Learning from Movies

Mahdi M. Kalayeh, Shervin Ardeshir, Lingyi Liu, Nagendra Kamath, Ashok Chandrashekar

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments arXiv admin note: substantial text overlap with arXiv:2106.08513

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.12366 2022-04-27 cs.MM 79%

Robust Audio-Visual Instance Discrimination via Active Contrastive Set Mining

Hanyu Xuan, Yihong Xu, Shuo Chen, Zhiliang Wu, Jian Yang, Yan Yan, Xavier Alameda-Pineda

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.MM

Comments 7 pages, 4 figures, accepted at IJCAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08686 2022-04-21 cs.SD eess.AS 79%

Audio-Visual Wake Word Spotting System For MISP Challenge 2021

Yanguang Xu, Jianwei Sun, Yang Han, Shuaijiang Zhao, Chaoyang Mei, Tingwei Guo, Shuran Zhou, Chuandong Xie, Wei Zou, Xiangang Li, Shuran Zhou, Chuandong Xie, Wei Zou, Xiangang Li

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 eess.AS

Comments Accepted to ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏