arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4585 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4585 篇

2501.09506 2025-01-22 cs.LG cs.SD eess.AS eess.IV 79%

Multimodal Marvels of Deep Learning in Medical Diagnosis: A Comprehensive Review of COVID-19 Detection

Md Shofiqul Islam, Khondokar Fida Hasan, Hasibul Hossain Shajeeb, Humayan Kabir Rana, Md Saifur Rahmand, Md Munirul Hasan, AKM Azad, Ibrahim Abdullah, Mohammad Ali Moni

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments 43 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08124 2025-01-15 eess.AS cs.SD 79%

Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech

Mareike Daeglau, Juergen Otten, Giso Grimm, Bojana Mirkovic, Volker Hohmann, Stefan Debener

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06530 2025-01-14 eess.AS cs.SD 79%

Multi-modal Speech Enhancement with Limited Electromyography Channels

Fuyuan Feng, Longting Xu, Rohan Kumar Das

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00481 2025-01-09 eess.AS cs.SD 79%

DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module

Xinyu Wang, Haotian Jiang, Haolin Huang, Yu Fang, Mengjie Xu, Qian Wang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted to ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04138 2025-01-09 cs.CL 79%

"Yeah Right!" -- Do LLMs Exhibit Multimodal Feature Transfer?

Benjamin Reichman, Kartik Talamadupula

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01401 2025-01-03 eess.AS 79%

VoiceVector: Multimodal Enrolment Vectors for Speaker Separation

Akam Rahimi, Triantafyllos Afouras, Andrew Zisserman

专题命中 音频语音多模态 :multimodal(title);audio-visual(abstract);分类 eess.AS

Journal ref 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00029 2025-01-03 cs.CL cs.IR cs.LG 79%

A Breadth-First Catalog of Text Processing, Speech Processing and Multimodal Research in South Asian Languages

Pranav Gupta

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20872 2025-01-03 cs.CV 79%

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing

Langyu Wang, Bingke Zhu, Yingying Chen, Jinqiao Wang

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by ICASSP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20467 2024-12-31 cs.CL 79%

Utilizing Multimodal Data for Edge Case Robust Call-sign Recognition and Understanding

Alexander Blatt, Dietrich Klakow

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19563 2024-12-30 cs.CV 79%

Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing

Yongbiao Gao, Xiangcheng Sun, Guohua Lv, Deng Yu, Sijiu Niu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16771 2024-12-24 cs.CV 79%

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization

Tan-Hanh Pham, Hoang-Nam Le, Phu-Vinh Nguyen, Chris Ngo, Truong-Son Hy

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13387 2024-12-19 eess.AS cs.SD 79%

Deep Speech Synthesis from Multimodal Articulatory Representations

Peter Wu, Bohan Yu, Kevin Scheck, Alan W Black, Aditi S. Krishnapriyan, Irene Y. Chen, Tanja Schultz, Shinji Watanabe, Gopala K. Anumanchipalli

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10103 2024-12-16 cs.CL 79%

AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation

Xiyuan Gao, Shubhi Bansal, Kushaan Gowda, Zhu Li, Shekhar Nayak, Nagendra Kumar, Matt Coler

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments This is a preprint version of the paper, submitted and under review at the IEEE Transactions on Affective Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09317 2024-12-13 cs.SD cs.AI cs.CV cs.MM eess.AS 79%

Multimodal Sentiment Analysis based on Video and Audio Inputs

Antonio Fernandez, Suzan Awinat

专题命中 音频语音多模态 :multimodal(title);分类 cs.CV、cs.AI、cs.MM

Comments Presented as a full paper in the 15th International Conference on Emerging Ubiquitous Systems and Pervasive Networks (EUSPN 2024) October 28-30, 2024, Leuven, Belgium

Journal ref Procedia Computer Science, Volume 251, 2024, Pages 41-48, ISSN 1877-0509

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08529 2024-12-12 cs.CL 79%

TECO: Improving Multimodal Intent Recognition with Text Enhancement through Commonsense Knowledge Extraction

Quynh-Mai Thi Nguyen, Lan-Nhi Thi Nguyen, Cam-Van Thi Nguyen

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted at PACLIC 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16213 2024-12-12 cs.CV 79%

SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context

Jungang Li, Sicheng Tao, Yibo Yan, Xiaojie Gu, Haodong Xu, Xu Zheng, Yuanhuiyi Lyu, Linfeng Zhang, Xuming Hu

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments The publication has some processing errors (language short-cuts in synthetic data are not avoided) that invalidate some of the conclusions

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05315 2024-12-10 cs.CL cs.CY 79%

Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor

Ashwin Baluja

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10234 2024-11-18 cs.HC cs.AI 79%

Generative AI in Multimodal User Interfaces: Trends, Challenges, and Cross-Platform Adaptability

J. Bieniek, M. Rahouti, D. C. Verma

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10880 2024-11-15 cs.CL 79%

Exploring the Potential of Multimodal LLM with Knowledge-Intensive Multimodal ASR

Minghan Wang, Yuxia Wang, Thuy-Trang Vu, Ehsan Shareghi, Gholamreza Haffari

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06754 2024-11-12 cs.LG cs.AI 79%

Scaling Law Hypothesis for Multimodal Model

Qingyun Sun, Zhen Guo, PIN AI Team

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05603 2024-11-11 cs.CV 79%

Efficient Audio-Visual Fusion for Video Classification

Mahrukh Awan, Asmar Nadeem, Armin Mustafa

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments CVMP Short Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22112 2024-10-30 cs.MM 79%

Multimodal Semantic Communication for Generative Audio-Driven Video Conferencing

Haonan Tong, Haopeng Li, Hongyang Du, Zhaohui Yang, Changchuan Yin, Dusit Niyato

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments accepted by IEEE Wireless Communications Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21170 2024-10-29 cs.CV 79%

Joint Audio-Visual Idling Vehicle Detection with Streamlined Input Dependencies

Xiwen Li, Rehman Mohammed, Tristalee Mangin, Surojit Saha, Ross T Whitaker, Kerry E. Kelly, Tolga Tasdizen

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20916 2024-10-29 cs.CL 79%

NeuGPT: Unified multi-modal Neural GPT

Yiqian Yang, Yiqun Duan, Hyejeong Jo, Qiang Zhang, Renjing Xu, Oiwi Parker Jones, Xuming Hu, Chin-teng Lin, Hui Xiong

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20116 2024-10-29 cs.HC cs.AI 79%

Estuary: A Framework For Building Multimodal Low-Latency Real-Time Socially Interactive Agents

Spencer Lin, Basem Rizk, Miru Jun, Andy Artze, Caitlin Sullivan, Sharon Mozgai, Scott Fisher

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments To be published in ACM Intelligent Virtual Agents (IVA) 2024 [DOI: 10.1145/3652988.3696198] [ACM ISBN: 979-8-4007-0625-7/24/09]

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18882 2024-10-25 cs.CL 79%

A Survey of Multimodal Sarcasm Detection

Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanojia, Yu Kong, Marcos Zampieri

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Published in the Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence Survey Track. Pages 8020-8028

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03650 2024-10-22 cs.MM 79%

Towards Multimodal Emotional Support Conversation Systems

Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14917 2024-10-21 cs.CL 79%

With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models

Tyler Loakman, Yucheng Li, Chenghua Lin

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted to EMNLP 2024 (Camera Ready)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13803 2024-10-18 cs.AI 79%

A Pattern to Align Them All: Integrating Different Modalities to Define Multi-Modal Entities

Gianluca Apriceno, Valentina Tamma, Tania Bailoni, Jacopo de Berardinis, Mauro Dragoni

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

Comments 20 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07757 2024-10-11 cs.CV 79%

MMHead: Towards Fine-grained Multi-modal 3D Facial Animation

Sijing Wu, Yunhao Li, Yichao Yan, Huiyu Duan, Ziwei Liu, Guangtao Zhai

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

Comments Accepted by ACMMM 2024. Project page: https://wsj-sjtu.github.io/MMHead/

详情

展开后加载摘要…

URL PDF HTML 收藏