arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2202.12257 2022-03-08 cs.SD cs.LG eess.AS 57%

A Perceptual Measure for Evaluating the Resynthesis of Automatic Music Transcriptions

Federico Simonetta, Federico Avanzini, Stavros Ntalampiras

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by Multimedia Tools and Applications (2022); supplementary materials are in the latex sources

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.01265 2022-03-03 cs.CV 57%

Self-supervised Transformer for Deepfake Detection

Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Weiming Zhang, Nenghai Yu

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13066 2022-03-01 eess.AS cs.SD 57%

Revisiting Over-Smoothness in Text to Speech

Yi Ren, Xu Tan, Tao Qin, Zhou Zhao, Tie-Yan Liu

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02939 2022-02-22 eess.AS eess.SP 57%

Unsupervised Audio-Caption Aligning Learns Correspondences between Individual Sound Events and Textual Phrases

Huang Xie, Okko Räsänen, Konstantinos Drossos, Tuomas Virtanen

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments Accepted at ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.09418 2022-02-11 eess.AS cs.IR cs.SD 57%

Audio Retrieval with Natural Language Queries: A Benchmark Study

A. Sophia Koepke, Andreea-Maria Oncescu, João F. Henriques, Zeynep Akata, Samuel Albanie

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments Submitted to Transactions on Multimedia. arXiv admin note: substantial text overlap with arXiv:2105.02192

Journal ref IEEE Transactions on Multimedia 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.01155 2022-02-03 cs.CL 57%

The slurk Interaction Server Framework: Better Data for Better Dialog Models

Jana Götze, Maike Paetzel-Prüsmann, Wencke Liermann, Tim Diekmann, David Schlangen

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments submitted to LREC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06367 2022-02-01 cs.CV cs.LG 57%

Voice-assisted Image Labelling for Endoscopic Ultrasound Classification using Neural Networks

Ester Bonmati, Yipeng Hu, Alexander Grimwood, Gavin J. Johnson, George Goodchild, Margaret G. Keane, Kurinchi Gurusamy, Brian Davidson, Matthew J. Clarkson, Stephen P. Pereira, Dean C. Barratt

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Submitted to IEEE TMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.07150 2022-01-25 eess.AS 57%

Selective Listening by Synchronizing Speech with Lips

Zexu Pan, Ruijie Tao, Chenglin Xu, Haizhou Li

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted by TASLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.02192 2022-01-07 cs.RO cs.AI cs.HC cs.SY eess.SY 57%

A wearable sensor vest for social humanoid robots with GPGPU, IoT, and modular software architecture

Mohsen Jafarzadeh, Stephen Brooks, Shimeng Yu, Balakrishnan Prabhakaran, Yonas Tadesse

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

Comments This is the preprint version. The final version is published in Robotics and Autonomous Systems, Volume 139, 2021, Page 103536, ISSN 0921-8890, https://doi.org/10.1016/j.robot.2020.103536

Journal ref Robotics and Autonomous Systems, vol 139, page 103536, year 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.05302 2022-01-05 cs.MM 57%

Semantics-Consistent Representation Learning for Remote Sensing Image-Voice Retrieval

Hailong Ning, Bin Zhao, Yuan Yuan

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.13353 2021-12-28 cs.SD cs.LG eess.AS 57%

Novel Hybrid DNN Approaches for Speaker Verification in Emotional and Stressful Talking Environments

Ismail Shahin, Ali Bou Nassif, Nawel Nemmour, Ashraf Elnagar, Adi Alhudhaif, Kemal Polat

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 23 pages, 13 figures

Journal ref Published in Neural Computing and Applications. Vol. 33, issue 23, June 2021, pp. 16033-16055

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00803 2021-12-03 cs.CL 57%

A Review of Web Infodemic Analysis and Detection Trends across Multi-modalities using Deep Neural Networks

Chahat Raj, Priyanka Meel

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

Comments 47 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.10592 2021-11-23 cs.SD cs.HC cs.LG eess.AS 57%

Deep Spoken Keyword Spotting: An Overview

Iván López-Espejo, Zheng-Hua Tan, John Hansen, Jesper Jensen

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.05098 2021-11-23 cs.CL 57%

Position-based Contributive Embeddings for Aspect-Based Sentiment Analysis

Zijian Zhang, Chenxin Zhang, Jiangfeng Li, Qinpei Zhao

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 5 pages, 3 figures. Submitted to ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.08567 2021-11-17 cs.CV 57%

Joint Learning of Visual-Audio Saliency Prediction and Sound Source Localization on Multi-face Videos

Minglang Qiao, Yufan Liu, Mai Xu, Xin Deng, Bing Li, Weiming Hu, Ali Borji

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments 21 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.13442 2021-11-17 eess.AS cs.SD 57%

Multi-channel Multi-frame ADL-MVDR for Target Speech Separation

Zhuohuang Zhang, Yong Xu, Meng Yu, Shi-Xiong Zhang, Lianwu Chen, Donald S. Williamson, Dong Yu

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted by IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP); Demos available at https://zzhang68.github.io/mcmf-adl-mvdr/

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15941 2021-11-01 cs.LG cs.SD eess.AS 57%

Personalized breath based biometric authentication with wearable multimodality

Manh-Ha Bui, Viet-Anh Tran, Cuong Pham

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 7 pages (2 columns), 5 tables, 7 figures, submitted to ACM Multimedia 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.13634 2021-10-25 eess.AS cs.SD 57%

Don't Separate, Learn to Remix: End-to-End Neural Remixing with Joint Optimization

Haici Yang, Shivani Firodiya, Nicholas J. Bryan, Minje Kim

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02360 2021-10-07 eess.AS cs.SD 57%

Neural Pitch-Shifting and Time-Stretching with Controllable LPCNet

Max Morrison, Zeyu Jin, Nicholas J. Bryan, Juan-Pablo Caceres, Bryan Pardo

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Submitted to ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.04865 2021-09-28 cs.SD cs.GR eess.AS 57%

Learning Acoustic Scattering Fields for Dynamic Interactive Sound Propagation

Zhenyu Tang, Hsien-Yu Meng, Dinesh Manocha

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Journal ref 2021 IEEE Virtual Reality and 3D User Interfaces (VR) (pp. 835-844)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07879 2021-08-19 cs.AR cs.AI cs.ET cs.LG 57%

Edge AI without Compromise: Efficient, Versatile and Accurate Neurocomputing in Resistive Random-Access Memory

Weier Wan, Rajkumar Kubendran, Clemens Schaefer, S. Burc Eryilmaz, Wenqiang Zhang, Dabin Wu, Stephen Deiss, Priyanka Raina, He Qian, Bin Gao, Siddharth Joshi, Huaqiang Wu, H. -S. Philip Wong, Gert Cauwenberghs

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 34 pages, 14 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06720 2021-08-17 cs.CV 57%

Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational Autoencoders

Jing Li, Di Kang, Wenjie Pei, Xuefei Zhe, Ying Zhang, Zhenyu He, Linchao Bao

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04316 2021-08-13 cs.HC cs.AI cs.SD 57%

Generating Music and Generative Art from Brain activity

Ricardo Andres Diaz Rincon

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.01246 2021-08-04 cs.RO cs.CV 57%

AcousticFusion: Fusing Sound Source Localization to Visual SLAM in Dynamic Environments

Tianwei Zhang, Huayan Zhang, Xiaofei Li, Junfeng Chen, Tin Lun Lam, Sethu Vijayakumar

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Accepted by IROS-2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.11548 2021-07-28 cs.SD cs.GR eess.AS 57%

Dynamic Portal Occlusion for Precomputed Interactive Sound Propagation

Nikunj Raghuvanshi

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 6 pages, 5 figures, planning to submit to IEEE TVCG Short papers at a future date

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.02192 2021-07-23 cs.IR cs.SD eess.AS 57%

Audio Retrieval with Natural Language Queries

Andreea-Maria Oncescu, A. Sophia Koepke, João F. Henriques, Zeynep Akata, Samuel Albanie

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments Accepted at INTERSPEECH 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.12733 2021-07-16 cs.SD eess.AS 57%

Learning Fine-Grained Cross Modality Excitement for Speech Emotion Recognition

Hang Li, Wenbiao Ding, Zhongqin Wu, Zitao Liu

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments The Interspeech Conference, 2021 (INTERSPEECH 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.14016 2021-06-29 cs.MM 57%

An Attention Self-supervised Contrastive Learning based Three-stage Model for Hand Shape Feature Representation in Cued Speech

Jianrong Wang, Nan Gu, Mei Yu, Xuewei Li, Qiang Fang, Li Liu

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.12790 2021-06-25 cs.CV 57%

Towards Automatic Speech to Sign Language Generation

Parul Kapoor, Rudrabha Mukhopadhyay, Sindhu B Hegde, Vinay Namboodiri, C V Jawahar

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments 5 pages(including references), 5 figures, Accepted in Interspeech 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.01691 2021-06-10 eess.AS 57%

A Study of Incorporating Articulatory Movement Information in Speech Enhancement

Yu-Wen Chen, Kuo-Hsuan Hung, Shang-Yi Chuang, Jonathan Sherman, Xugang Lu, Yu Tsao

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏