arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2210.12701 2022-10-25 eess.AS cs.SD 57%

Speaker Identification from emotional and noisy speech data using learned voice segregation and Speech VGG

Shibani Hamsa, Ismail Shahin, Youssef Iraqi, Ernesto Damiani, Naoufel Werghi

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07091 2022-08-16 cs.SD cs.LG eess.AS 57%

Analysis of impact of emotions on target speech extraction and speech separation

Ján Švec, Kateřina Žmolíková, Martin Kocour, Marc Delcroix, Tsubasa Ochiai, Ladislav Mošner, Jan Černocký

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted to IWAENC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02246 2022-08-04 cs.LG cs.AI stat.ML 57%

AdaCat: Adaptive Categorical Discretization for Autoregressive Models

Qiyang Li, Ajay Jain, Pieter Abbeel

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Uncertainty in Artificial Intelligence (UAI) 2022 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.03074 2022-07-22 cs.SD eess.AS eess.IV 57%

Visual-Assisted Sound Source Depth Estimation in the Wild

Wei Sun, Lili Qiu

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 13 pages;in submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10326 2022-07-22 eess.AS cs.SD 57%

Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion

Zongyang Du, Berrak Sisman, Kun Zhou, Haizhou Li

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments Accepted by Interspeech 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07693 2022-07-19 cs.HC cs.AI cs.LG 57%

Towards Understanding Confusion and Affective States Under Communication Failures in Voice-Based Human-Machine Interaction

Sujeong Kim, Abhinav Garlapati, Jonah Lubin, Amir Tamrakar, Ajay Divakaran

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

Journal ref 2021 9th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.12910 2022-06-30 cs.LG cs.CV 57%

Sparse Centroid-Encoder: A Nonlinear Model for Feature Selection

Tomojit Ghosh, Michael Kirby

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments 13 pages,56 figures, 5 tables. Used 12 data sets and 5 state-of-the-art models for comparison

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.09121 2022-06-30 cs.LG cs.CV 57%

Single-Layer Vision Transformers for More Accurate Early Exits with Less Overhead

Arian Bakhtiarnia, Qi Zhang, Alexandros Iosifidis

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by Neural Networks journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.03112 2022-06-22 cs.LG cs.SD eess.AS 57%

Singapore Soundscape Site Selection Survey (S5): Identification of Characteristic Soundscapes of Singapore via Weighted k-means Clustering

Kenneth Ooi, Bhan Lam, Joo Young Hong, Karn N. Watcharasupat, Zhen-Ting Ong, Woon-Seng Gan

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 23 pages, 8 figures. Submitted to Sustainability

Journal ref MDPI Sustainability. 2022; 14(12):7485

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16843 2022-06-22 eess.AS cs.SD 57%

A Hybrid Continuity Loss to Reduce Over-Suppression for Time-domain Target Speaker Extraction

Zexu Pan, Meng Ge, Haizhou Li

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted by Interspeech2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09634 2022-06-20 cs.SD cs.LG eess.AS 57%

Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering

Samuel Lipping, Parthasaarathy Sudarsanam, Konstantinos Drossos, Tuomas Virtanen

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.15060 2022-06-15 cs.CL 57%

Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems

Ting-En Lin, Yuchuan Wu, Fei Huang, Luo Si, Jian Sun, Yongbin Li

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted by KDD 2022, ADS track

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04769 2022-06-13 cs.SD eess.AS 57%

CLAP: Learning Audio Concepts From Natural Language Supervision

Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, Huaming Wang

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04527 2022-05-24 eess.AS 57%

Transformer Network for Semantically-Aware and Speech-Driven Upper-Face Generation

Mireille Fares, Catherine Pelachaud, Nicolas Obin

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10266 2022-05-23 cs.CV 57%

Analysis of Co-Laughter Gesture Relationship on RGB videos in Dyadic Conversation Contex

Hugo Bohy, Ahmad Hammoudeh, Antoine Maiorca, Stéphane Dupont, Thierry Dutoit

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 5 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03684 2022-05-10 cs.MM 57%

Timestamp-independent Haptic-Visual Synchronization

Yiwen Xu, Liangtao Huang, Tiesong Zhao, Liqun Lin, Ying Fang

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.00514 2022-05-06 cs.CL 57%

Visualization: the missing factor in Simultaneous Speech Translation

Sara Papi, Matteo Negri, Marco Turchi

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL

Comments Accepted at CLIC-it 2021

Journal ref Italian Conference on Computational Linguistics 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00129 2022-05-03 cs.LG cs.CV 57%

Gaze-enhanced Crossmodal Embeddings for Emotion Recognition

Ahmed Abdou, Ekta Sood, Philipp Müller, Andreas Bulling

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.07537 2022-05-02 eess.AS cs.LG cs.SD 57%

Toward Degradation-Robust Voice Conversion

Chien-yu Huang, Kai-Wei Chang, Hung-yi Lee

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments To appear in the proceedings of ICASSP 2022, equal contribution from first two authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.11356 2022-04-26 cs.LG cs.CL cs.SI 57%

Hate Me Not: Detecting Hate Inducing Memes in Code Switched Languages

Kshitij Rajput, Raghav Kapoor, Kaushal Rai, Preeti Kaur

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments To be published in 2022 Americas Conference on Information Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08451 2022-04-19 cs.CV 57%

Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion

Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li, Trevor Darrell, Angjoo Kanazawa, Shiry Ginosar

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05738 2022-04-13 eess.AS cs.SD 57%

Text-Driven Separation of Arbitrary Sounds

Kevin Kilgour, Beat Gfeller, Qingqing Huang, Aren Jansen, Scott Wisdom, Marco Tagliasacchi

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Submitted to INTERSPEECH 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.01725 2022-04-06 cs.CV 57%

Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading

Minsu Kim, Jeong Hun Yeo, Yong Man Ro

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Published at AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00088 2022-04-04 cs.SD cs.LG eess.AS q-bio.QM 57%

Speech and the n-Back task as a lens into depression. How combining both may allow us to isolate different core symptoms of depression

Salvatore Fara, Stefano Goria, Emilia Molimpakis, Nicholas Cummins

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Submitted to Interspeech 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16497 2022-04-04 cs.HC cs.AI 57%

An Artificial Intelligence Browser Architecture (AIBA) For Our Kind and Others: A Voice Name System Speech implementation with two warrants, Wake Neutrality and Value Preservation of Personally Identifiable Information

Brian Subirana

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16937 2022-04-01 cs.SD cs.LG eess.AS 57%

HiFi-VC: High Quality ASR-Based Voice Conversion

A. Kashkin, I. Karpukhin, S. Shishkin

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments Submitted to INTERSPEECH 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15253 2022-03-30 cs.SD cs.LG eess.AS 57%

NeuraGen-A Low-Resource Neural Network based approach for Gender Classification

Shankhanil Ghosh, Chhanda Saha, Naagamani Molakathaala

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.13412 2022-03-28 cs.CV 57%

Self-Supervised Predictive Learning: A Negative-Free Method for Sound Source Localization in Visual Scenes

Zengjie Song, Yuxi Wang, Junsong Fan, Tieniu Tan, Zhaoxiang Zhang

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Camera-ready, CVPR 2022. Code: https://github.com/zjsong/SSPL

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.11443 2022-03-23 cs.CL 57%

Demo of the Linguistic Field Data Management and Analysis System -- LiFE

Siddharth Singh, Ritesh Kumar, Shyam Ratan, Sonal Sinha

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL

Comments Accepted in the 19th International Conference on Natural Language Processing (ICON-2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.11368 2022-03-23 cs.CV 57%

Audio visual character profiles for detecting background characters in entertainment media

Rahul Sharma, Shrikanth Narayanan

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments submitted to ICIP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏