arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2204.00088 2022-04-04 cs.SD cs.LG eess.AS q-bio.QM 57%

Speech and the n-Back task as a lens into depression. How combining both may allow us to isolate different core symptoms of depression

Salvatore Fara, Stefano Goria, Emilia Molimpakis, Nicholas Cummins

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Submitted to Interspeech 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16497 2022-04-04 cs.HC cs.AI 57%

An Artificial Intelligence Browser Architecture (AIBA) For Our Kind and Others: A Voice Name System Speech implementation with two warrants, Wake Neutrality and Value Preservation of Personally Identifiable Information

Brian Subirana

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16937 2022-04-01 cs.SD cs.LG eess.AS 57%

HiFi-VC: High Quality ASR-Based Voice Conversion

A. Kashkin, I. Karpukhin, S. Shishkin

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments Submitted to INTERSPEECH 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15253 2022-03-30 cs.SD cs.LG eess.AS 57%

NeuraGen-A Low-Resource Neural Network based approach for Gender Classification

Shankhanil Ghosh, Chhanda Saha, Naagamani Molakathaala

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13412 2022-03-28 cs.CV 57%

Self-Supervised Predictive Learning: A Negative-Free Method for Sound Source Localization in Visual Scenes

Zengjie Song, Yuxi Wang, Junsong Fan, Tieniu Tan, Zhaoxiang Zhang

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Camera-ready, CVPR 2022. Code: https://github.com/zjsong/SSPL

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.11443 2022-03-23 cs.CL 57%

Demo of the Linguistic Field Data Management and Analysis System -- LiFE

Siddharth Singh, Ritesh Kumar, Shyam Ratan, Sonal Sinha

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL

Comments Accepted in the 19th International Conference on Natural Language Processing (ICON-2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.11368 2022-03-23 cs.CV 57%

Audio visual character profiles for detecting background characters in entertainment media

Rahul Sharma, Shrikanth Narayanan

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments submitted to ICIP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.12257 2022-03-08 cs.SD cs.LG eess.AS 57%

A Perceptual Measure for Evaluating the Resynthesis of Automatic Music Transcriptions

Federico Simonetta, Federico Avanzini, Stavros Ntalampiras

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by Multimedia Tools and Applications (2022); supplementary materials are in the latex sources

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.01265 2022-03-03 cs.CV 57%

Self-supervised Transformer for Deepfake Detection

Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Weiming Zhang, Nenghai Yu

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13066 2022-03-01 eess.AS cs.SD 57%

Revisiting Over-Smoothness in Text to Speech

Yi Ren, Xu Tan, Tao Qin, Zhou Zhao, Tie-Yan Liu

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by ACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02939 2022-02-22 eess.AS eess.SP 57%

Unsupervised Audio-Caption Aligning Learns Correspondences between Individual Sound Events and Textual Phrases

Huang Xie, Okko Räsänen, Konstantinos Drossos, Tuomas Virtanen

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments Accepted at ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.09418 2022-02-11 eess.AS cs.IR cs.SD 57%

Audio Retrieval with Natural Language Queries: A Benchmark Study

A. Sophia Koepke, Andreea-Maria Oncescu, João F. Henriques, Zeynep Akata, Samuel Albanie

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments Submitted to Transactions on Multimedia. arXiv admin note: substantial text overlap with arXiv:2105.02192

Journal ref IEEE Transactions on Multimedia 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.01155 2022-02-03 cs.CL 57%

The slurk Interaction Server Framework: Better Data for Better Dialog Models

Jana Götze, Maike Paetzel-Prüsmann, Wencke Liermann, Tim Diekmann, David Schlangen

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments submitted to LREC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.06367 2022-02-01 cs.CV cs.LG 57%

Voice-assisted Image Labelling for Endoscopic Ultrasound Classification using Neural Networks

Ester Bonmati, Yipeng Hu, Alexander Grimwood, Gavin J. Johnson, George Goodchild, Margaret G. Keane, Kurinchi Gurusamy, Brian Davidson, Matthew J. Clarkson, Stephen P. Pereira, Dean C. Barratt

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Submitted to IEEE TMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.07150 2022-01-25 eess.AS 57%

Selective Listening by Synchronizing Speech with Lips

Zexu Pan, Ruijie Tao, Chenglin Xu, Haizhou Li

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted by TASLP

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.02192 2022-01-07 cs.RO cs.AI cs.HC cs.SY eess.SY 57%

A wearable sensor vest for social humanoid robots with GPGPU, IoT, and modular software architecture

Mohsen Jafarzadeh, Stephen Brooks, Shimeng Yu, Balakrishnan Prabhakaran, Yonas Tadesse

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

Comments This is the preprint version. The final version is published in Robotics and Autonomous Systems, Volume 139, 2021, Page 103536, ISSN 0921-8890, https://doi.org/10.1016/j.robot.2020.103536

Journal ref Robotics and Autonomous Systems, vol 139, page 103536, year 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.05302 2022-01-05 cs.MM 57%

Semantics-Consistent Representation Learning for Remote Sensing Image-Voice Retrieval

Hailong Ning, Bin Zhao, Yuan Yuan

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.13353 2021-12-28 cs.SD cs.LG eess.AS 57%

Novel Hybrid DNN Approaches for Speaker Verification in Emotional and Stressful Talking Environments

Ismail Shahin, Ali Bou Nassif, Nawel Nemmour, Ashraf Elnagar, Adi Alhudhaif, Kemal Polat

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 23 pages, 13 figures

Journal ref Published in Neural Computing and Applications. Vol. 33, issue 23, June 2021, pp. 16033-16055

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.00803 2021-12-03 cs.CL 57%

A Review of Web Infodemic Analysis and Detection Trends across Multi-modalities using Deep Neural Networks

Chahat Raj, Priyanka Meel

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

Comments 47 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.10592 2021-11-23 cs.SD cs.HC cs.LG eess.AS 57%

Deep Spoken Keyword Spotting: An Overview

Iván López-Espejo, Zheng-Hua Tan, John Hansen, Jesper Jensen

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.05098 2021-11-23 cs.CL 57%

Position-based Contributive Embeddings for Aspect-Based Sentiment Analysis

Zijian Zhang, Chenxin Zhang, Jiangfeng Li, Qinpei Zhao

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 5 pages, 3 figures. Submitted to ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.08567 2021-11-17 cs.CV 57%

Joint Learning of Visual-Audio Saliency Prediction and Sound Source Localization on Multi-face Videos

Minglang Qiao, Yufan Liu, Mai Xu, Xin Deng, Bing Li, Weiming Hu, Ali Borji

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments 21 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.13442 2021-11-17 eess.AS cs.SD 57%

Multi-channel Multi-frame ADL-MVDR for Target Speech Separation

Zhuohuang Zhang, Yong Xu, Meng Yu, Shi-Xiong Zhang, Lianwu Chen, Donald S. Williamson, Dong Yu

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted by IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP); Demos available at https://zzhang68.github.io/mcmf-adl-mvdr/

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.15941 2021-11-01 cs.LG cs.SD eess.AS 57%

Personalized breath based biometric authentication with wearable multimodality

Manh-Ha Bui, Viet-Anh Tran, Cuong Pham

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 7 pages (2 columns), 5 tables, 7 figures, submitted to ACM Multimedia 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.13634 2021-10-25 eess.AS cs.SD 57%

Don't Separate, Learn to Remix: End-to-End Neural Remixing with Joint Optimization

Haici Yang, Shivani Firodiya, Nicholas J. Bryan, Minje Kim

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02360 2021-10-07 eess.AS cs.SD 57%

Neural Pitch-Shifting and Time-Stretching with Controllable LPCNet

Max Morrison, Zeyu Jin, Nicholas J. Bryan, Juan-Pablo Caceres, Bryan Pardo

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Submitted to ICASSP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.04865 2021-09-28 cs.SD cs.GR eess.AS 57%

Learning Acoustic Scattering Fields for Dynamic Interactive Sound Propagation

Zhenyu Tang, Hsien-Yu Meng, Dinesh Manocha

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Journal ref 2021 IEEE Virtual Reality and 3D User Interfaces (VR) (pp. 835-844)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07879 2021-08-19 cs.AR cs.AI cs.ET cs.LG 57%

Edge AI without Compromise: Efficient, Versatile and Accurate Neurocomputing in Resistive Random-Access Memory

Weier Wan, Rajkumar Kubendran, Clemens Schaefer, S. Burc Eryilmaz, Wenqiang Zhang, Dabin Wu, Stephen Deiss, Priyanka Raina, He Qian, Bin Gao, Siddharth Joshi, Huaqiang Wu, H. -S. Philip Wong, Gert Cauwenberghs

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 34 pages, 14 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.06720 2021-08-17 cs.CV 57%

Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational Autoencoders

Jing Li, Di Kang, Wenjie Pei, Xuefei Zhe, Ying Zhang, Zhenyu He, Linchao Bao

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.04316 2021-08-13 cs.HC cs.AI cs.SD 57%

Generating Music and Generative Art from Brain activity

Ricardo Andres Diaz Rincon

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏