arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2404.02098 2024-04-03 cs.CV 57%

BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition

Alexandros Haliassos, Andreas Zinonos, Rodrigo Mira, Stavros Petridis, Maja Pantic

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments ICASSP 2024. Code: https://github.com/ahaliassos/raven

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07676 2024-04-02 cs.CR cs.CL cs.LG 57%

Composite Backdoor Attacks Against Large Language Models

Hai Huang, Zhengyu Zhao, Michael Backes, Yun Shen, Yang Zhang

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments To Appear in Findings of the Association for Computational Linguistics: NAACL 2024, June 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16078 2024-03-26 cs.SD eess.AS 57%

Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy

Wenxuan Wu, Xueyuan Chen, Xixin Wu, Haizhou Li, Helen Meng

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted by IJCNN 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05730 2024-03-15 cs.MM 57%

AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild

Yongkang Yin, Xu Li, Ying Shan, Yuexian Zou

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.14363 2024-03-13 cs.AI 57%

Mobile Foundation Model as Firmware

Jinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang, Xin Yuan, Zeling Zhang, Xiang Li, Dingge Zhang, Hanzi Mei, Xianqing Jia, Shangguang Wang, Mengwei Xu

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 17 pages, 15 figures, published to ACM MobiCom'24

Journal ref The 30th Annual International Conference on Mobile Computing and Networking, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06282 2024-03-12 cs.LG cs.CV cs.IR 57%

MuseChat: A Conversational Music Recommendation System for Videos

Zhikang Dong, Bin Chen, Xiulong Liu, Pawel Polak, Peng Zhang

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04594 2024-03-08 cs.SD eess.AS 57%

A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds

Xuenan Xu, Xiaohang Xu, Zeyu Xie, Pingyue Zhang, Mengyue Wu, Kai Yu

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.07902 2024-03-06 cs.SD eess.AS 57%

BLAT: Bootstrapping Language-Audio Pre-training based on AudioSet Tag-guided Synthetic Data

Xuenan Xu, Zhiling Zhang, Zelin Zhou, Pingyue Zhang, Zeyu Xie, Mengyue Wu, Kenny Q. Zhu

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.02112 2024-03-05 cs.CV 57%

A New Perspective on Smiling and Laughter Detection: Intensity Levels Matter

Hugo Bohy, Kevin El Haddad, Thierry Dutoit

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Journal ref In 2022 10th International Conference on Affective Computing and Intelligent Interaction (ACII) (pp. 1-8). IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.01670 2024-03-05 eess.AS cs.SD 57%

6DoF SELD: Sound Event Localization and Detection Using Microphones and Motion Tracking Sensors on self-motioning human

Masahiro Yasuda, Shoichiro Saito, Akira Nakayama, Noboru Harada

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments ICASSP2024 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07300 2024-02-28 cs.HC cs.MM 57%

SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers

Zheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, Toby Jia-Jun Li

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13004 2024-02-21 cs.CV 57%

Comparison of Conventional Hybrid and CTC/Attention Decoders for Continuous Visual Speech Recognition

David Gimeno-Gómez, Carlos-D. Martínez-Hinarejos

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Accepted at the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10790 2024-02-21 eess.AS cs.SD 57%

Listen, Think, and Understand

Yuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky, James Glass

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted at ICLR 2024. Code, dataset, and models are available at https://github.com/YuanGongND/ltu. The interactive demo is at https://huggingface.co/spaces/yuangongfdu/ltu

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05767 2024-02-08 cs.SD eess.AS 57%

Natural Language Supervision for General-Purpose Audio Representations

Benjamin Elizalde, Soham Deshmukh, Huaming Wang

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02899 2024-02-06 eess.AS 57%

Positive and negative sampling strategies for self-supervised learning on audio-video data

Shanshan Wang, Soumya Tripathy, Toni Heittola, Annamaria Mesaros

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02830 2024-02-06 eess.AS 57%

Automatic Detection of Depression in Speech Using Ensemble Convolutional Neural Networks

Adrián Vázquez-Romero, Ascensión Gallardo-Antolín

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Journal ref A. Vazquez-Romero and A. Gallardo-Antolin Entropy 22, no. 6: 688 (2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01385 2024-02-05 eess.AS cs.SD 57%

Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen

María Sánchez, Laura Fernández, Julián Arias, Mateo Cámara, Giulia Comini, Adam Gabrys, José Luis Blanco, Juan Ignacio Godino, Luis Alfonso Hernández

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 10 pages, in Spanish, Tecniacústica

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10946 2024-01-23 cs.HC cs.AI 57%

Self context-aware emotion perception on human-robot interaction

Zihan Lin, Francisco Cruz, Eduardo Benitez Sandoval

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Australasian Conference on Robotics and Automation (ACRA). 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10751 2024-01-22 cs.AI cs.CY cs.SC 57%

EFO: the Emotion Frame Ontology

Stefano De Giorgis, Aldo Gangemi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.12883 2024-01-17 cs.HC cs.AI 57%

Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video

Matthew Groh, Aruna Sankaranarayanan, Nikhil Singh, Dong Young Kim, Andrew Lippman, Rosalind Picard

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.05757 2024-01-12 cs.SD eess.AS physics.class-ph 57%

Intuitive Control of Scraping and Rubbing Through Audio-tactile Synthesis

Mitsuko Aramaki, Corentin Bernard, Richard Kronland-Martinet, Samuel Poirot, Sølvi Ystad

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Journal ref 16th International Symposium on CMMR, Future University Hakodate, Japan; Asia University, Japan; Nihon University, Japan; Laboratoire PRISM, Marseille, France, Nov 2023, Tokyo, France

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04919 2024-01-09 cs.SD eess.AS 57%

Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion

Binzhu Sha, Xu Li, Zhiyong Wu, Ying Shan, Helen Meng

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.14314 2024-01-03 cs.SD eess.AS 57%

Adapter Incremental Continual Learning of Efficient Audio Spectrogram Transformers

Nithish Muthuchamy Selvaraj, Xiaobao Guo, Adams Kong, Bingquan Shen, Alex Kot

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.06382 2024-01-02 cs.SD cs.LG eess.AS 57%

Phoneme Hallucinator: One-shot Voice Conversion via Set Expansion

Siyuan Shan, Yang Li, Amartya Banerjee, Junier B. Oliva

专题命中 音频语音多模态 :any-to-any(abstract);分类 eess.AS

Comments AAAI 2024 Demo, Codes: https://phonemehallucinator.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14021 2023-12-22 eess.AS cs.LG cs.SD eess.IV eess.SP 57%

Leveraging Visual Supervision for Array-based Active Speaker Detection and Localization

Davide Berghi, Philip J. B. Jackson

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13873 2023-12-22 cs.SD eess.AS 57%

Self-Supervised Adaptive AV Fusion Module for Pre-Trained ASR Models

Christopher Simic, Tobias Bocklet

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted at ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06034 2023-12-12 cs.AI 57%

Modeling Uncertainty in Personalized Emotion Prediction with Normalizing Flows

Piotr Miłkowski, Konrad Karanowski, Patryk Wielopolski, Jan Kocoń, Przemysław Kazienko, Maciej Zięba

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 10 pages, 8 figures, SENTIRE'23 (ICDM 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05709 2023-11-13 cs.CV cs.LG 57%

OmniVec: Learning robust representations with cross modal sharing

Siddharth Srivastava, Gaurav Sharma

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV

Comments Accepted to WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.01458 2023-11-03 cs.CV cs.LG 57%

Detecting Deepfakes Without Seeing Any

Tal Reiss, Bar Cavia, Yedid Hoshen

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Our code is available at https://github.com/talreiss/FACTOR

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.19559 2023-11-03 cs.CV 57%

Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning

Changsheng Lv, Shuai Zhang, Yapeng Tian, Mengshi Qi, Huadong Ma

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments To be published in 37th Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏