arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4585 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4585 篇

2403.06354 2024-03-12 cs.CL 79%

Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages

Michael Andersland

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04866 2024-03-11 cs.AI 79%

A Modular End-to-End Multimodal Learning Method for Structured and Unstructured Data

Marco D Alessandro, Enrique Calabrés, Mikel Elkano

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments 8 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16124 2024-02-27 cs.CV 79%

AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation

Yasheng Sun, Wenqing Chu, Hang Zhou, Kaisiyuan Wang, Hideki Koike

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14579 2024-02-26 cs.CV 79%

Real-Time Idling Vehicles Detection using Combined Audio-Visual Deep Learning

Xiwen Li, Tristalee Mangin, Surojit Saha, Evan Blanchard, Dillon Tang, Henry Poppe, Nathan Searle, Ouk Choi, Kerry Kelly, Ross Whitaker

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13851 2024-02-22 cs.CV 79%

VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models

Jiawei Liang, Siyuan Liang, Man Luo, Aishan Liu, Dongchen Han, Ee-Chien Chang, Xiaochun Cao

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11297 2024-02-20 cs.CL 79%

MMMModal -- Multi-Images Multi-Audio Multi-turn Multi-Modal

Husein Zolkepli, Aisyah Razak, Kamarul Adha, Ariff Nazhan

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01967 2024-02-20 cs.CL 79%

MasonPerplexity at Multimodal Hate Speech Event Detection 2024: Hate Speech and Target Detection Using Transformer Ensembles

Amrita Ganguly, Al Nahian Bin Emran, Sadiya Sayara Chowdhury Puspo, Md Nishat Raihan, Dhiman Goswami, Marcos Zampieri

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09245 2024-02-15 eess.AS cs.LG eess.SP 79%

Overview of the L3DAS23 Challenge on Audio-Visual Extended Reality

Christian Marinoni, Riccardo Fosco Gramaccioni, Changan Chen, Aurelio Uncini, Danilo Comminiello

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted to 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07327 2024-02-13 cs.AI 79%

Multi-Modal Emotion Recognition by Text, Speech and Video Using Pretrained Transformers

Minoo Shayaninasab, Bagher Babaali

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14378 2024-02-12 cs.LG cs.SD eess.AS 79%

Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification

Anirudh S. Sundar, Chao-Han Huck Yang, David M. Chan, Shalini Ghosh, Venkatesh Ravichandran, Phani Sankar Nidadavolu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments 5 pages, 1 figure, ICASSP 2024 Workshop on Self-supervision in Audio, Speech and Beyond

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.09138 2024-02-07 cs.CV 79%

An Open-source Benchmark of Deep Learning Models for Audio-visual Apparent and Self-reported Personality Recognition

Rongfan Liao, Siyang Song, Hatice Gunes

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Affective Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.18084 2024-02-01 cs.CV cs.RO 79%

Binding Touch to Everything: Learning Unified Multimodal Tactile Representations

Fengyu Yang, Chao Feng, Ziyang Chen, Hyoungseob Park, Daniel Wang, Yiming Dou, Ziyao Zeng, Xien Chen, Rit Gangopadhyay, Andrew Owens, Alex Wong

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.10687 2024-02-01 eess.AS cs.SD 79%

MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis

Wenhao Guan, Yishuang Li, Tao Li, Hukai Huang, Feng Wang, Jiayan Lin, Lingyan Huang, Lin Li, Qingyang Hong

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments Accepted at AAAI2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.05669 2024-01-17 cs.CV 79%

Multi-Modal Gaze Following in Conversational Scenarios

Yuqi Hou, Zhongqun Zhang, Nora Horanyi, Jaewon Moon, Yihua Cheng, Hyung Jin Chang

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02746 2024-01-08 cs.CV 79%

Reading Between the Frames: Multi-Modal Depression Detection in Videos from Non-Verbal Cues

David Gimeno-Gómez, Ana-Maria Bucur, Adrian Cosma, Carlos-David Martínez-Hinarejos, Paolo Rosso

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted at 46th European Conference on Information Retrieval (ECIR 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00430 2024-01-04 cs.AI 79%

Brain-Conditional Multimodal Synthesis: A Survey and Taxonomy

Weijian Mai, Jian Zhang, Pengfei Fang, Zhijun Zhang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.00424 2024-01-02 cs.CL 79%

SDIF-DA: A Shallow-to-Deep Interaction Framework with Data Augmentation for Multi-modal Intent Detection

Shijue Huang, Libo Qin, Bingbing Wang, Geng Tu, Ruifeng Xu

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted by ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17262 2024-01-01 cs.CL cs.LG 79%

Multimodal Classification of Teaching Activities from University Lecture Recordings

Oscar Sapena, Eva Onaindia

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 18 pages

Journal ref Appl. Sci. 2022, 12, 4785

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08288 2023-12-20 cs.CV 79%

Improving Audio-Visual Segmentation with Bidirectional Generation

Dawei Hao, Yuxin Mao, Bowen He, Xiaodong Han, Yuchao Dai, Yiran Zhong

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments AAAI Camera Ready. Dawei Hao and Yuxin Mao contribute equality to this paper. Yiran Zhong is the corresponding author. The code will be released at https://github.com/OpenNLPLab/AVS-bidirectional

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12558 2023-12-19 cs.CV 79%

Hyperbolic Audio-visual Zero-shot Learning

Jie Hong, Zeeshan Hayder, Junlin Han, Pengfei Fang, Mehrtash Harandi, Lars Petersson

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments ICCV 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08850 2023-12-15 cs.SD eess.AS 79%

Hourglass-AVSR: Down-Up Sampling-based Computational Efficiency Model for Audio-Visual Speech Recognition

Fan Yu, Haoxu Wang, Ziyang Ma, Shiliang Zhang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted by ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.11871 2023-12-14 cs.CL 79%

Cem Mil Podcasts: A Spoken Portuguese Document Corpus For Multi-modal, Multi-lingual and Multi-Dialect Information Access Research

Ekaterina Garmash, Edgar Tanaka, Ann Clifton, Joana Correia, Sharmistha Jat, Winstead Zhu, Rosie Jones, Jussi Karlgren

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

Comments 12 pages, 1 figure

Journal ref Volume 14163 of Lecture Notes in Computer Science, pages 48-59, Springer, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.04131 2023-12-08 eess.AS cs.SD 79%

Joint Training or Not: An Exploration of Pre-trained Speech Models in Audio-Visual Speaker Diarization

Huan Zhao, Li Zhang, Yue Li, Yannan Wang, Hongji Wang, Wei Rao, Qing Wang, Lei Xie

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.03632 2023-12-07 cs.SD cs.LG eess.AS 79%

Multimodal Data and Resource Efficient Device-Directed Speech Detection with Large Foundation Models

Dominik Wagner, Alexander Churchill, Siddharth Sigtia, Panayiotis Georgiou, Matt Mirsamadi, Aarshee Mishra, Erik Marchi

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01568 2023-12-05 cs.HC cs.SD eess.AS 79%

Multimodal Speech Emotion Recognition Using Modality-specific Self-Supervised Frameworks

Rutherford Agbeshi Patamia, Paulo E. Santos, Kingsley Nketia Acheampong, Favour Ekong, Kwabena Sarpong, She Kun

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17177 2023-11-30 cs.CV 79%

THInImg: Cross-modal Steganography for Presenting Talking Heads in Images

Lin Zhao, Hongxuan Li, Xuefei Ning, Xinru Jiang

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at WACV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16471 2023-11-29 cs.CV 79%

A Unified Framework for Multimodal, Multi-Part Human Motion Synthesis

Zixiang Zhou, Yu Wan, Baoyuan Wang

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments 19 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16446 2023-11-29 cs.CV 79%

Centre Stage: Centricity-based Audio-Visual Temporal Action Detection

Hanyuan Wang, Majid Mirmehdi, Dima Damen, Toby Perrett

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted to VUA workshop at BMVC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07860 2023-11-29 cs.SD cs.LG eess.AS 79%

EPG2S: Speech Generation and Speech Enhancement based on Electropalatography and Audio Signals using Multimodal Learning

Li-Chin Chen, Po-Hsun Chen, Richard Tzong-Han Tsai, Yu Tsao

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Accepted By IEEE Signal Processing Letter

Journal ref IEEE Signal Processing Letters, vol. 29, p. 2582-2586, 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13165 2023-11-23 cs.AI 79%

Multimodal Large Language Models: A Survey

Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan, Philip S. Yu

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments IEEE BigData 2023. 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏