arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2403.02112 2024-03-05 cs.CV 57%

A New Perspective on Smiling and Laughter Detection: Intensity Levels Matter

Hugo Bohy, Kevin El Haddad, Thierry Dutoit

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Journal ref In 2022 10th International Conference on Affective Computing and Intelligent Interaction (ACII) (pp. 1-8). IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01670 2024-03-05 eess.AS cs.SD 57%

6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on self-motioning human

Masahiro Yasuda, Shoichiro Saito, Akira Nakayama, Noboru Harada

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments ICASSP2024 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07300 2024-02-28 cs.HC cs.MM 57%

SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers

Zheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, Toby Jia-Jun Li

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13004 2024-02-21 cs.CV 57%

Comparison of Conventional Hybrid and CTC/Attention Decoders for Continuous Visual Speech Recognition

David Gimeno-Gómez, Carlos-D. Martínez-Hinarejos

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Accepted at the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10790 2024-02-21 eess.AS cs.SD 57%

Listen, Think, and Understand

Yuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky, James Glass

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted at ICLR 2024. Code, dataset, and models are available at https://github.com/YuanGongND/ltu. The interactive demo is at https://huggingface.co/spaces/yuangongfdu/ltu

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05767 2024-02-08 cs.SD eess.AS 57%

Natural Language Supervision for General-Purpose Audio Representations

Benjamin Elizalde, Soham Deshmukh, Huaming Wang

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02899 2024-02-06 eess.AS 57%

Positive and negative sampling strategies for self-supervised learning on audio-video data

Shanshan Wang, Soumya Tripathy, Toni Heittola, Annamaria Mesaros

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02830 2024-02-06 eess.AS 57%

Automatic Detection of Depression in Speech Using Ensemble Convolutional Neural Networks

Adrián Vázquez-Romero, Ascensión Gallardo-Antolín

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Journal ref A. Vazquez-Romero and A. Gallardo-Antolin Entropy 22, no. 6: 688 (2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01385 2024-02-05 eess.AS cs.SD 57%

Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen

María Sánchez, Laura Fernández, Julián Arias, Mateo Cámara, Giulia Comini, Adam Gabrys, José Luis Blanco, Juan Ignacio Godino, Luis Alfonso Hernández

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 10 pages, in Spanish, Tecniacústica

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10946 2024-01-23 cs.HC cs.AI 57%

Self context-aware emotion perception on human-robot interaction

Zihan Lin, Francisco Cruz, Eduardo Benitez Sandoval

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Australasian Conference on Robotics and Automation (ACRA). 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10751 2024-01-22 cs.AI cs.CY cs.SC 57%

EFO: the Emotion Frame Ontology

Stefano De Giorgis, Aldo Gangemi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.12883 2024-01-17 cs.HC cs.AI 57%

Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video

Matthew Groh, Aruna Sankaranarayanan, Nikhil Singh, Dong Young Kim, Andrew Lippman, Rosalind Picard

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05757 2024-01-12 cs.SD eess.AS physics.class-ph 57%

Intuitive Control of Scraping and Rubbing Through Audio-tactile Synthesis

Mitsuko Aramaki, Corentin Bernard, Richard Kronland-Martinet, Samuel Poirot, Sølvi Ystad

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Journal ref 16th International Symposium on CMMR, Future University Hakodate, Japan; Asia University, Japan; Nihon University, Japan; Laboratoire PRISM, Marseille, France, Nov 2023, Tokyo, France

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04919 2024-01-09 cs.SD eess.AS 57%

Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion

Binzhu Sha, Xu Li, Zhiyong Wu, Ying Shan, Helen Meng

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14314 2024-01-03 cs.SD eess.AS 57%

Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers

Nithish Muthuchamy Selvaraj, Xiaobao Guo, Adams Kong, Bingquan Shen, Alex Kot

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06382 2024-01-02 cs.SD cs.LG eess.AS 57%

Phoneme Hallucinator: One-shot Voice Conversion via Set Expansion

Siyuan Shan, Yang Li, Amartya Banerjee, Junier B. Oliva

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments AAAI 2024 Demo, Codes: https://phonemehallucinator.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14021 2023-12-22 eess.AS cs.LG cs.SD eess.IV eess.SP 57%

Leveraging Visual Supervision for Array-based Active Speaker Detection and Localization

Davide Berghi, Philip J. B. Jackson

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13873 2023-12-22 cs.SD eess.AS 57%

Self-Supervised Adaptive AV Fusion Module for Pre-Trained ASR Models

Christopher Simic, Tobias Bocklet

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted at ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06034 2023-12-12 cs.AI 57%

Modeling Uncertainty in Personalized Emotion Prediction with Normalizing Flows

Piotr Miłkowski, Konrad Karanowski, Patryk Wielopolski, Jan Kocoń, Przemysław Kazienko, Maciej Zięba

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 10 pages, 8 figures, SENTIRE'23 (ICDM 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05709 2023-11-13 cs.CV cs.LG 57%

OmniVec: Learning robust representations with cross modal sharing

Siddharth Srivastava, Gaurav Sharma

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01458 2023-11-03 cs.CV cs.LG 57%

Detecting Deepfakes Without Seeing Any

Tal Reiss, Bar Cavia, Yedid Hoshen

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Our code is available at https://github.com/talreiss/FACTOR

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19559 2023-11-03 cs.CV 57%

Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning

Changsheng Lv, Shuai Zhang, Yapeng Tian, Mengshi Qi, Huadong Ma

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments To be published in 37th Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05374 2023-10-19 physics.med-ph cs.SD eess.AS eess.IV 57%

Towards Ultrasound Tongue Image prediction from EEG during speech production

Tamás Gábor Csapó, Frigyes Viktor Arthur, Péter Nagy, Ádám Boncz

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments accepted at Interspeech 2023

Journal ref Proceedings of Interspeech 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11295 2023-10-18 cs.CV cs.CG 57%

CorrTalk: Correlation Between Hierarchical Speech and Facial Activity Variances for 3D Animation

Zhaojie Chu, Kailing Guo, Xiaofen Xing, Yilin Lan, Bolun Cai, Xiangmin Xu

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10179 2023-10-17 eess.AS cs.SD 57%

Advancing Audio Emotion and Intent Recognition with Large Pre-Trained Models and Bayesian Inference

Dejan Porjazovski, Yaroslav Getman, Tamás Grósz, Mikko Kurimo

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted at ACMM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.08749 2023-10-11 cs.CL cs.LG 57%

Utilizing Longitudinal Chest X-Rays and Reports to Pre-Fill Radiology Reports

Qingqing Zhu, Tejas Sudharshan Mathai, Pritam Mukherjee, Yifan Peng, Ronald M. Summers, Zhiyong Lu

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05304 2023-10-10 cs.CV 57%

GestSync: Determining who is speaking without a talking head

Sindhu B Hegde, Andrew Zisserman

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Accepted in BMVC 2023, 10 pages paper, 7 pages supplementary, 7 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.11096 2023-10-09 cs.SD cs.LG eess.AS 57%

Robust One-Shot Singing Voice Conversion

Naoya Takahashi, Mayank Kumar Singh, Yuki Mitsufuji

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.02419 2023-10-03 cs.CV 57%

TM2D: Bimodality Driven 3D Dance Generation via Music-Text Integration

Kehong Gong, Dongze Lian, Heng Chang, Chuan Guo, Zihang Jiang, Xinxin Zuo, Michael Bi Mi, Xinchao Wang

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted by ICCV2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12111 2023-09-22 cs.SD cs.IR cs.LG eess.AS 57%

Passage Summarization with Recurrent Models for Audio-Sheet Music Retrieval

Luis Carvalho, Gerhard Widmer

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments In Proceedings of the 24th Conference of the International Society for Music Information Retrieval (ISMIR 2023), Milan, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏