arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2211.09376 2022-11-18 cs.SD cs.LG eess.AS 57%

Balanced Deep CCA for Bird Vocalization Detection

Sumit Kumar, B. Anshuman, Linus Ruettimann, Richard H. R. Hahnloser, Vipul Arora

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.10608 2022-11-18 cs.CL 57%

Dodging the Data Bottleneck: Automatic Subtitling with Automatically Segmented ST Corpora

Sara Papi, Alina Karakanta, Matteo Negri, Marco Turchi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Journal ref AACL 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.14099 2022-11-18 cs.CL 57%

Climate and Weather: Inspecting Depression Detection via Emotion Recognition

Wen Wu, Mengyue Wu, Kai Yu

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Journal ref ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6262-6266

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.05446 2022-11-11 cs.SD cs.CR cs.LG eess.AS 57%

Privacy-Utility Balanced Voice De-Identification Using Adversarial Examples

Meng Chen, Li Lu, Jiadi Yu, Yingying Chen, Zhongjie Ba, Feng Lin, Kui Ren

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03511 2022-11-08 cs.CL 57%

End-to-End Evaluation of a Spoken Dialogue System for Learning Basic Mathematics

Eda Okur, Saurav Sahay, Roddy Fuentes Alba, Lama Nachman

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Proceedings of the 1st Workshop on Mathematical Natural Language Processing (MathNLP) at EMNLP 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.03019 2022-11-08 cs.CV 57%

Hear The Flow: Optical Flow-Based Self-Supervised Visual Sound Source Localization

Dennis Fedorishin, Deen Dayal Mohan, Bhavin Jawade, Srirangaraj Setlur, Venu Govindaraju

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Accepted to WACV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.02336 2022-11-07 cs.SD eess.AS 57%

Improving Speech Prosody of Audiobook Text-to-Speech Synthesis with Acoustic and Textual Contexts

Detai Xin, Sharath Adavanne, Federico Ang, Ashish Kulkarni, Shinnosuke Takamichi, Hiroshi Saruwatari

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.12701 2022-10-25 eess.AS cs.SD 57%

Speaker Identification from emotional and noisy speech data using learned voice segregation and Speech VGG

Shibani Hamsa, Ismail Shahin, Youssef Iraqi, Ernesto Damiani, Naoufel Werghi

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07091 2022-08-16 cs.SD cs.LG eess.AS 57%

Analysis of impact of emotions on target speech extraction and speech separation

Ján Švec, Kateřina Žmolíková, Martin Kocour, Marc Delcroix, Tsubasa Ochiai, Ladislav Mošner, Jan Černocký

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted to IWAENC 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.02246 2022-08-04 cs.LG cs.AI stat.ML 57%

AdaCat: Adaptive Categorical Discretization for Autoregressive Models

Qiyang Li, Ajay Jain, Pieter Abbeel

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Uncertainty in Artificial Intelligence (UAI) 2022 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.03074 2022-07-22 cs.SD eess.AS eess.IV 57%

Visual-Assisted Sound Source Depth Estimation in the Wild

Wei Sun, Lili Qiu

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 13 pages;in submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.10326 2022-07-22 eess.AS cs.SD 57%

Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion

Zongyang Du, Berrak Sisman, Kun Zhou, Haizhou Li

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments Accepted by Interspeech 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.07693 2022-07-19 cs.HC cs.AI cs.LG 57%

Towards Understanding Confusion and Affective States Under Communication Failures in Voice-Based Human-Machine Interaction

Sujeong Kim, Abhinav Garlapati, Jonah Lubin, Amir Tamrakar, Ajay Divakaran

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

Journal ref 2021 9th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.12910 2022-06-30 cs.LG cs.CV 57%

Sparse Centroid-Encoder: A Nonlinear Model for Feature Selection

Tomojit Ghosh, Michael Kirby

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments 13 pages,56 figures, 5 tables. Used 12 data sets and 5 state-of-the-art models for comparison

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.09121 2022-06-30 cs.LG cs.CV 57%

Single-Layer Vision Transformers for More Accurate Early Exits with Less Overhead

Arian Bakhtiarnia, Qi Zhang, Alexandros Iosifidis

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Accepted by Neural Networks journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.03112 2022-06-22 cs.LG cs.SD eess.AS 57%

Singapore Soundscape Site Selection Survey (S5): Identification of Characteristic Soundscapes of Singapore via Weighted k-means Clustering

Kenneth Ooi, Bhan Lam, Joo Young Hong, Karn N. Watcharasupat, Zhen-Ting Ong, Woon-Seng Gan

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 23 pages, 8 figures. Submitted to Sustainability

Journal ref MDPI Sustainability. 2022; 14(12):7485

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.16843 2022-06-22 eess.AS cs.SD 57%

A Hybrid Continuity Loss to Reduce Over-Suppression for Time-domain Target Speaker Extraction

Zexu Pan, Meng Ge, Haizhou Li

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted by Interspeech2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.09634 2022-06-20 cs.SD cs.LG eess.AS 57%

Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering

Samuel Lipping, Parthasaarathy Sudarsanam, Konstantinos Drossos, Tuomas Virtanen

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.15060 2022-06-15 cs.CL 57%

Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems

Ting-En Lin, Yuchuan Wu, Fei Huang, Luo Si, Jian Sun, Yongbin Li

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted by KDD 2022, ADS track

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04769 2022-06-13 cs.SD eess.AS 57%

CLAP: Learning Audio Concepts From Natural Language Supervision

Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, Huaming Wang

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04527 2022-05-24 eess.AS 57%

Transformer Network for Semantically-Aware and Speech-Driven Upper-Face Generation

Mireille Fares, Catherine Pelachaud, Nicolas Obin

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.10266 2022-05-23 cs.CV 57%

Analysis of Co-Laughter Gesture Relationship on RGB videos in Dyadic Conversation Contex

Hugo Bohy, Ahmad Hammoudeh, Antoine Maiorca, Stéphane Dupont, Thierry Dutoit

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 5 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03684 2022-05-10 cs.MM 57%

Timestamp-independent Haptic-Visual Synchronization

Yiwen Xu, Liangtao Huang, Tiesong Zhao, Liqun Lin, Ying Fang

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.00514 2022-05-06 cs.CL 57%

Visualization: the missing factor in Simultaneous Speech Translation

Sara Papi, Matteo Negri, Marco Turchi

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL

Comments Accepted at CLIC-it 2021

Journal ref Italian Conference on Computational Linguistics 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00129 2022-05-03 cs.LG cs.CV 57%

Gaze-enhanced Crossmodal Embeddings for Emotion Recognition

Ahmed Abdou, Ekta Sood, Philipp Müller, Andreas Bulling

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.07537 2022-05-02 eess.AS cs.LG cs.SD 57%

Toward Degradation-Robust Voice Conversion

Chien-yu Huang, Kai-Wei Chang, Hung-yi Lee

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments To appear in the proceedings of ICASSP 2022, equal contribution from first two authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.11356 2022-04-26 cs.LG cs.CL cs.SI 57%

Hate Me Not: Detecting Hate Inducing Memes in Code Switched Languages

Kshitij Rajput, Raghav Kapoor, Kaushal Rai, Preeti Kaur

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments To be published in 2022 Americas Conference on Information Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08451 2022-04-19 cs.CV 57%

Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion

Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li, Trevor Darrell, Angjoo Kanazawa, Shiry Ginosar

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05738 2022-04-13 eess.AS cs.SD 57%

Text-Driven Separation of Arbitrary Sounds

Kevin Kilgour, Beat Gfeller, Qingqing Huang, Aren Jansen, Scott Wisdom, Marco Tagliasacchi

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Submitted to INTERSPEECH 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.01725 2022-04-06 cs.CV 57%

Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading

Minsu Kim, Jeong Hun Yeo, Yong Man Ro

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Published at AAAI 2022

详情

展开后加载摘要…

URL PDF HTML 收藏