arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4585 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4585 篇

2509.20741 2025-09-26 eess.AS cs.ET cs.LG 79%

Real-Time System for Audio-Visual Target Speech Enhancement

T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang

机构 * Bose Corporation(博世公司) Georgia Institute of Technology(佐治亚理工学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted into WASPAA 2025 demo session

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11976 2025-09-24 cs.SD eess.AS 79%

PoolingVQ: A VQVAE Variant for Reducing Audio Redundancy and Boosting Multi-Modal Fusion in Music Emotion Analysis

Dinghao Zou, Yicheng Gong, Xiaokang Li, Xin Cao, Sunbowen Lee

机构 * Wuhan University of Science and Technology(武汉科技大学)

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02864 2025-09-23 cs.RO cs.CV 79%

The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio

Renhao Wang, Haoran Geng, Tingle Li, Feishi Wang, Gopala Anumanchipalli, Trevor Darrell, Boyi Li, Pieter Abbeel, Jitendra Malik, Alexei A. Efros

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments Conference on Robot Learning 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16023 2025-09-22 eess.AS 79%

Interpreting the Role of Visemes in Audio-Visual Speech Recognition

Aristeidis Papadopoulos, Naomi Harte

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted into Automatic Speech Recognition and Understanding- ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14891 2025-09-19 cs.MM cs.IR cs.SD 79%

Music4All A+A: A Multimodal Dataset for Music Information Retrieval Tasks

Jonas Geiger, Marta Moscati, Shah Nawaz, Markus Schedl

机构 * Johannes Kepler University Linz(约翰内斯·开普勒大学林茨) Human-centered AI Group, AI Lab, Linz Institute of Technology(以人为本的人工智能小组、人工智能实验室、林茨技术研究所)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments 7 pages, 6 tables, IEEE International Conference on Content-Based Multimedia Indexing (IEEE CBMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14379 2025-09-19 eess.AS cs.LG 79%

Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior

Yochai Yemini, Rami Ben-Ari, Sharon Gannot, Ethan Fetaya

机构 * Faculty of Engineering, Bar-Ilan University(巴伊兰大学工程学院) OriginAI

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11937 2025-09-16 cs.SE cs.AI 79%

MMORE: Massive Multimodal Open RAG & Extraction

Alexandre Sallinen, Stefan Krsteski, Paul Teiletche, Marc-Antoine Allard, Baptiste Lecoeur, Michael Zhang, Fabrice Nemo, David Kalajdzic, Matthias Meyer, Mary-Anne Hartley

机构 * École Polytechnique Fédérale de Lausanne (EPFL), Switzerland(瑞士联邦理工学院洛桑校区) ETH Zürich, Switzerland(瑞士苏黎世联邦理工学院) T.H. Chan School of Public Health, Harvard University, USA(哈佛大学T.H. Chan公共卫生学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments This paper was originally submitted to the CODEML workshop for ICML 2025. 9 pages (including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11183 2025-09-16 cs.SD eess.AS 79%

WeaveMuse: An Open Agentic System for Multimodal Music Understanding and Generation

Emmanouil Karystinaios

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Accepted at Large Language Models for Music & Audio Workshop (LLM4MA) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18039 2025-09-16 cs.AI 79%

MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind

Zheng Zhang, Nuoqian Xiao, Qi Chai, Deheng Ye, Hao Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tencent(腾讯)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09859 2025-09-15 cs.CV cs.LG 79%

WAVE-DETR Multi-Modal Visible and Acoustic Real-Life Drone Detector

Razvan Stefanescu, Ethan Oh, Ruben Vazquez, Chris Mesterharm, Constantin Serban, Ritu Chadha

机构 * Peraton Labs(珀顿实验室)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 11 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20511 2025-09-10 cs.CL 79%

Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospects

Chengyan Wu, Yiqiang Cai, Yang Liu, Pengxu Zhu, Yun Xue, Ziwei Gong, Julia Hirschberg, Bolei Ma

机构 * Guangdong Provincial Key Laboratory of Quantum Engineering and Quantum Materials(广东省量子工程与量子材料重点实验室) School of Electronic Science and Engineering (School of Microelectronics), South China Normal University(华南师范大学电子科学学院(微电子学院)) North Carolina Central University(北卡罗来纳中央大学) Georgia Institute of Technology(佐治亚理工学院) Columbia University(哥伦比亚大学) LMU Munich & Munich Center for Machine Learning(慕尼黑大学及慕尼黑机器学习中心)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06074 2025-09-09 cs.CL 79%

Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis

Zhenqi Jia, Rui Liu, Berrak Sisman, Haizhou Li

机构 * Inner Mongolia University(内蒙古大学) Center for Language and Speech Processing (CLSP)(语言与语音处理中心) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)人工智能学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05205 2025-09-08 eess.AS cs.SD 79%

MEAN-RIR: Multi-Modal Environment-Aware Network for Robust Room Impulse Response Estimation

Jiajian Chen, Jiakang Chen, Hang Chen, Qing Wang, Yu Gao, Jun Du

机构 * University of Science and Technology of China(科学技术大学) AI Research Center, Midea Group (Shanghai) Co.,Ltd.(美的集团(上海)有限公司人工智能研究中心)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments Accepted by ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04605 2025-09-08 cs.CL 79%

Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects

Xiyuan Gao, Shekhar Nayak, Matt Coler

机构 * Campus Fryslân, University of Groningen(格罗宁根大学弗里桑校区)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments 20 pages, 7 figures, Submitted to IEEE Transactions on Affective Computing

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04215 2025-09-05 cs.SD cs.IR cs.MM 79%

PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music

Hayeon Bang, Eunjin Choi, Seungheon Doh, Juhan Nam

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments Accepted for publication at the 26th International Society for Music Information Retrieval Conference (ISMIR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03084 2025-09-03 cs.LG cs.AI 79%

Adversarial Attacks in Multimodal Systems: A Practitioner's Survey

Shashank Kapoor, Sanjay Surendranath Girija, Lakshit Arora, Dipen Pradhan, Ankit Shetgaonkar, Aman Raj

机构 * Google(谷歌)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments Accepted in IEEE COMPSAC 2025

Journal ref 2025 IEEE 49th Annual Computers, Software, and Applications Conference (COMPSAC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20513 2025-08-29 cs.SD cs.MM 79%

MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening

Yongqi Shao, Binxin Mei, Cong Tan, Hong Huo, Tao Fang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20221 2025-08-29 cs.CV 79%

Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos

Mert Cokelek, Halit Ozsoy, Nevrez Imamoglu, Cagri Ozcinar, Inci Ayhan, Erkut Erdem, Aykut Erdem

机构 * Department of Computer Science and Engineering, Koç University(计算机科学与工程系,科克大学) Department of Psychology, Boğaziçi University(心理学系,博多伊大学) National Institute of Advanced Industrial Science and Technology (AIST), Intelligent Platforms Research Institute(国家先进工业科学与技术研究院(AIST),智能平台研究机构) Department of Psychology, Boğaziçi University University(心理学系,博多伊大学) Department of Computer Engineering, Hacettepe University(计算机工程系,哈切塞特佩大学) KUIS AI Center(KUIS人工智能中心)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted for publication in IEEE Transaction on Pattern Analysis and Machine Intelligence (IEEE TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19483 2025-08-28 eess.AS 79%

Audio-Visual Feature Synchronization for Robust Speech Enhancement in Hearing Aids

Nasir Saleem, Mandar Gogate, Kia Dashtipour, Adeel Hussain, Usman Anwar, Adewale Adetomi, Tughrul Arslan, Amir Hussain

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Preprint of the paper presented at Euronoise 2025 Malaga, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16606 2025-08-26 cs.HC cs.AI cs.LG 79%

Multimodal Appearance based Gaze-Controlled Virtual Keyboard with Synchronous Asynchronous Interaction for Low-Resource Settings

Yogesh Kumar Meena, Manish Salvi

机构 * Human-AI Interaction (HAIx) Lab, IIT Gandhinagar(人机交互(HAIx)实验室,印度加尔各答理工学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16143 2025-08-25 cs.RO cs.AI 79%

Take That for Me: Multimodal Exophora Resolution with Interactive Questioning for Ambiguous Out-of-View Instructions

Akira Oyama, Shoichi Hasegawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi

机构 * Ritsumeikan University(立命馆大学) Soka University(早稻田大学) Kyoto University(京都大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments See website at https://emergentsystemlabstudent.github.io/MIEL/. Accepted at IEEE RO-MAN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12227 2025-08-22 cs.CL 79%

Arabic Multimodal Machine Learning: Datasets, Applications, Approaches, and Challenges

Abdelhamid Haouhat, Slimane Bellaouar, Attia Nehar, Hadda Cherroun, Ahmed Abdelali

机构 * Ziane Achour University(赞赞·阿赫尔大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11362 2025-08-18 cs.SD eess.AS 79%

Mitigating Category Imbalance: Fosafer System for the Multimodal Emotion and Intent Joint Understanding Challenge

Honghong Wang, Yankai Wang, Dejun Zhang, Jing Deng, Rong Zheng

机构 * Beijing Fosafer Information Technology Co., Ltd.(北京福萨弗信息科技有限公司)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments 2 pages. pubilshed by ICASSP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09913 2025-08-14 cs.CV 79%

SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection

Yachao Liang, Min Yu, Gang Li, Jianguo Jiang, Boquan Li, Feng Yu, Ning Zhang, Xiang Meng, Weiqing Huang

机构 * Institute of Information Engineering(信息工程研究所) Chinese Academy of Sciences(中国科学院) School of Cyber Security University of Chinese Academy of Sciences(中国科学院网络安全学院) Deakin University(德肯大学) Harbin Engineering University(哈尔滨工程大学) Institute of Computing Technology Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of Forensic Science Ministry of Public Security(公安部刑事科学技术研究所)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by NeurIPS 2024

Journal ref Advances in Neural Information Processing Systems, Volume 37, Pages 86124-86144, Year 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09702 2025-08-14 eess.AS cs.SD 79%

$\text{M}^3\text{PDB}$: A Multimodal, Multi-Label, Multilingual Prompt Database for Speech Generation

Boyu Zhu, Cheng Gong, Muyang Wu, Ruihao Jing, Fan Liu, Xiaolei Zhang, Chi Zhang, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAl), China Telecom(人工智能研究院(TeleAl),中国电信) School of Marine Science and Technology, Northwestern Polytechnical University(海洋科学与技术学院,西北工业大学)

专题命中 音频语音多模态 :multimodal(title);multi-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08042 2025-08-12 cs.IR cs.AI 79%

Multi-modal Adaptive Mixture of Experts for Cold-start Recommendation

Van-Khang Nguyen, Duc-Hoang Pham, Huy-Son Nguyen, Cam-Van Thi Nguyen, Hoang-Quynh Le, Duc-Trong Le

机构 * VNU University of Engineering and Technology(越南工程大学) Delft University of Technology(代尔夫特理工大学)

专题命中 音频语音多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02516 2025-08-12 cs.CV 79%

Engagement Prediction of Short Videos with Large Multimodal Models

Wei Sun, Linhan Cao, Yuqin Cao, Weixia Zhang, Wen Wen, Kaiwei Zhang, Zijian Chen, Fangfang Lu, Xiongkuo Min, Guangtao Zhai

机构 * East China Normal University(华东师范大学) Shanghai Jiao Tong University(上海交通大学) City University of Hong Kong(香港城市大学) Shanghai University of Electric Power(上海电力大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments The proposed method achieves first place in the ICCV VQualA 2025 EVQA-SnapUGC Challenge on short-form video engagement prediction

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07384 2025-08-07 cs.SD eess.AS 79%

AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment

Yu Chen, Hongxu Zhu, Jiadong Wang, Kainan Chen, Xinyuan Qian

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21448 2025-08-05 eess.AS cs.ET cs.LG 79%

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang

机构 * School of Music(音乐学院) Research(研究)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted into Interspeech 2025; corrected author name typo

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00784 2025-08-04 cs.AI 79%

Unraveling Hidden Representations: A Multi-Modal Layer Analysis for Better Synthetic Content Forensics

Tom Or, Omri Azencot

机构 * Ben Gurion University of the Negev(本· Gurion 内盖夫大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏