arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4585 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4585 篇

2507.23544 2025-08-01 cs.RO cs.CV cs.HC 79%

User Experience Estimation in Human-Robot Interaction Via Multi-Instance Learning of Multimodal Social Signals

Ryo Miyoshi, Yuki Okafuji, Takuya Iwamoto, Junya Nakanishi, Jun Baba

机构 * CyberAgent Osaka University(大阪大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments This paper has been accepted for presentation at IEEE/RSJ International Conference on Intelligent Robots and Systems 2025 (IROS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08052 2025-08-01 eess.AS 79%

Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization

Ming Cheng, Ming Li

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Accepted by IEEE Transactions on Audio, Speech, and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20579 2025-07-29 cs.CV 79%

AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations

Zhixi Cai, Kartik Kuckreja, Shreya Ghosh, Akanksha Chuchra, Muhammad Haris Khan, Usman Tariq, Tom Gedeon, Abhinav Dhall

机构 * Monash University(墨尔本大学) Curtin University(Curtin大学) American University of Sharjah(沙迦美国大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16736 2025-07-23 cs.CV 79%

DFR: A Decompose-Fuse-Reconstruct Framework for Multi-Modal Few-Shot Segmentation

Shuai Chen, Fanman Meng, Xiwei Zhang, Haoran Wei, Chenhao Wu, Qingbo Wu, Hongliang Li

机构 * School of Information and Communication Engineering(信息与通信工程学院) University of Electronic Science and Technology of China(电子科学与技术大学)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05328 2025-07-23 cs.CV 79%

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs

Lidong Lu, Guo Chen, Zhiqi Li, Yicheng Liu, Tong Lu

机构 * Nanjing University(南京大学)

专题命中 音频语音多模态 :audio-visual(title);multimodal(abstract);分类 cs.CV

Comments 21 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15072 2025-07-22 cs.HC cs.AI 79%

NavVI: A Telerobotic Simulation with Multimodal Feedback for Visually Impaired Navigation in Warehouse Environments

Maisha Maimuna, Minhaz Bin Farukee, Sama Nikanfar, Mahfuza Siddiqua, Ayon Roy, Fillia Makedon

机构 * Department of Computer Science and Engineering(计算机科学与工程系) The University of Texas at Arlington(德克萨斯理工大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09116 2025-07-22 cs.SD eess.AS 79%

Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition

Bingshen Mu, Kun Wei, Pengcheng Guo, Lei Xie

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments IEEE Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12972 2025-07-18 eess.AS cs.SD 79%

AVFSNet: Audio-Visual Speech Separation for Flexible Number of Speakers with Multi-Scale and Multi-Task Learning

Daning Zhang, Ying Wei

机构 * School of Control Science and Engineering, Shandong University(控制科学与工程学院,山东大学) Industrial Technology Research Institte of Shandong Province(山东省工业技术研究院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06566 2025-07-10 eess.AS 79%

Training Strategies for Modality Dropout Resilient Multi-Modal Target Speaker Extraction

Srikanth Korse, Mohamed Elminshawi, Emanuel A. P. Habets, Srikanth Raj Chetupalli

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments Published in ICASSPW 2024 (HSCMA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.02732 2025-07-10 cs.SD cs.LG eess.AS 79%

Multimodal Lyrics-Rhythm Matching

Callie C. Liao, Duoduo Liao, Jesse Guessford

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Accepted by 2022 IEEE International Conference on Big Data (IEEE Big Data 2022)

Journal ref IEEE BigData, Year: 2022; Pages: 3622-3630

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23623 2025-07-01 cs.CV 79%

Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

Shaofei Huang, Rui Ling, Tianrui Hui, Hongyu Li, Xu Zhou, Shifeng Zhang, Si Liu, Richang Hong, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) Chinese Academy of Sciences(中国科学院) Beihang University(北航) Sangfor Technologies(深信服技术)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by CVPR 2025; Code: https://github.com/spyflying/VCT_AVS; Models: https://huggingface.co/nowherespyfly/VCT_AVS

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23271 2025-07-01 cs.CV 79%

Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation

Jinxing Zhou, Zhihui Li, Yongqiang Yu, Yanghao Zhou, Ruohao Guo, Guangyao Li, Yuxin Mao, Mingfei Han, Xiaojun Chang, Meng Wang

机构 * MBZUAI Hefei University of Technology(合肥工业大学) University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Peking University(北京大学) Tsinghua University(清华大学) OpenNLP Lab(OpenNLP实验室)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22926 2025-07-01 cs.HC cs.GR cs.MM 79%

Coordinated 2D-3D Visualization of Volumetric Medical Data in XR with Multimodal Interactions

Qixuan Liu, Shi Qiu, Yinqiao Wang, Xiwen Wu, Kenneth Siu Ho Chok, Chi-Wing Fu, Pheng-Ann Heng

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.MM

Comments IEEE VIS 2025 Short Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.24066 2025-07-01 eess.AS eess.SP 79%

Cough-E: A multimodal, privacy-preserving cough detection algorithm for the edge

Stefano Albini, Lara Orlandic, Jonathan Dan, Jérôme Thevenot, Tomas Teijeiro, Denisa Andreea Constantinescu, David Atienza

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20945 2025-06-27 cs.SD eess.AS 79%

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis

Rui Niu, Weihao Wu, Jie Chen, Long Ma, Zhiyong Wu

机构 * Shenzhen International Graduate School, Tsinghua University, Shenzhen, China(深圳国际研究生院,清华大学,深圳,中国)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 eess.AS

Comments Accepted by ICME2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19603 2025-06-25 cs.CL cs.SI 79%

Social Hatred: Efficient Multimodal Detection of Hatemongers

Tom Marzea, Abraham Israeli, Oren Tsur

机构 * Ben Gurion University(本古里安大学) University of Michigan(密歇根大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

Comments To be published in WOAH, July 2025. arXiv admin note: text overlap with arXiv:2409.14464

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13419 2025-06-17 eess.IV cs.CV 79%

Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos

Riku Takahashi, Ryugo Morita, Jinjia Zhou

机构 * Hosei University(恒生大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted to ICMR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10331 2025-06-13 cs.CV eess.IV 79%

Research on Audio-Visual Quality Assessment Dataset and Method for User-Generated Omnidirectional Video

Fei Zhao, Da Pan, Zelu Qi, Ping Shi

机构 * School of Information and Communication Engineering(信息与通信工程学院)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Our paper has been accepted by ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06759 2025-06-10 cs.CV 79%

LitMAS: A Lightweight and Generalized Multi-Modal Anti-Spoofing Framework for Biometric Security

Nidheesh Gorthi, Kartik Thakral, Rishabh Ranjan, Richa Singh, Mayank Vatsa

机构 * Indian Institute of Information Technology Kottayam(印度信息技术学院科塔亚姆) Indian Institute of Technology Jodhpur(印度理工学院朱罗普尔)

专题命中 音频语音多模态 :multi-modal(title);cross-modal(abstract);分类 cs.CV

Comments Accepted in Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03980 2025-06-05 cs.CL 79%

Voice Activity Projection Model with Multimodal Encoders

Takeshi Saga, Catherine Pelachaud

机构 * Sorbonne University(索邦大学) CNRS(国家科学研究中心)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02470 2025-06-04 cs.AI 79%

A Smart Multimodal Healthcare Copilot with Powerful LLM Reasoning

Xuejiao Zhao, Siyan Liu, Su-Yin Yang, Chunyan Miao

机构 * Joint NTU-UBC Research Centre of Excellence in Active Living for the Elderly (LILY), NTU(联合NTU-UBC老龄化积极生活卓越研究中心(LILY),NTU) College of Computing and Data Science, Nanyang Technological University (NTU), Singapore(计算与数据科学学院,南洋理工大学(NTU),新加坡) Tan Tock Seng Hospital, Singapore(坦 tok sing 医院,新加坡) Woodlands Health, Singapore(伍德兰兹健康,新加坡)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02178 2025-06-04 cs.SD cs.CL 79%

Cocktail-Party Audio-Visual Speech Recognition

Thai-Binh Nguyen, Ngoc-Quan Pham, Alexander Waibel

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CL

Comments Accepted at Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01270 2025-06-03 eess.AS cs.SD 79%

Online Audio-Visual Autoregressive Speaker Extraction

Zexu Pan, Wupeng Wang, Shengkui Zhao, Chong Zhang, Kun Zhou, Yukun Ma, Bin Ma

机构 * Alibaba Group(阿里巴巴集团) Singapore National University of Singapore(新加坡国立大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS

Comments Interspeech2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07217 2025-06-03 cs.SD cs.CV 79%

ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation

Zixuan Wang, Chi-Keung Tang, Yu-Wing Tai

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Dartmouth College(达特茅斯学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

Comments Project page: https://vincent2311.github.io/ReelWave_demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23298 2025-05-30 cs.SD cs.IR eess.AS 79%

Bridging the Gap Between Semantic and User Preference Spaces for Multi-modal Music Representation Learning

Xiaofeng Pan, Jing Chen, Haitong Zhang, Menglin Xing, Jiayi Wei, Xuefeng Mu, Zhongqian Xie

机构 * NetEase Inc.(网易公司)

专题命中 音频语音多模态 :multi-modal(title,abstract);分类 eess.AS

Comments ICMR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10034 2025-05-30 cs.AI 79%

The First MPDD Challenge: Multimodal Personality-aware Depression Detection

Changzeng Fu, Zelin Fu, Qi Zhang, Xinhe Kuang, Jiacheng Dong, Kaifeng Su, Yikai Su, Wenbo Shi, Junfeng Yao, Yuliang Zhao, Shiqi Zhao, Jiadong Wang, Siyang Song, Chaoran Liu, Yuichiro Yoshikawa, Björn Schuller, Hiroshi Ishiguro

机构 * Northeastern University(东北大学) University of Technology Sydney(悉尼大学) Xiamen University(厦门大学) Technical University of Munich(慕尼黑技术大学) University of Cambridge(剑桥大学) National Information Institute(国家信息研究所) Osaka University(大阪大学) Imperial College London(伦敦帝国理工学院)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

Comments This paper has been accepted as part of the MPDD Challenge in the ACMMM 2025 Grand Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19938 2025-05-27 cs.CV 79%

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning

Wenrui Li, Penghong Wang, Xingtao Wang, Wangmeng Zuo, Xiaopeng Fan, Yonghong Tian

机构 * Harbin Institute of Technology(哈尔滨工业大学) Harbin Institute of Technology Suzhou Research Institute(哈尔滨工业大学苏州研究院) Peking University(北京大学) School of AI for Science(科学人工智能学院) Peng Cheng Laboratory(鹏城实验室)

专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV

Comments Accepted by IEEE TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11450 2025-05-27 cs.CV 79%

VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition

Puneet Kumar, Sarthak Malik, Balasubramanian Raman, Xiaobai Li

机构 * Center for Machine Vision and Signal Analysis, University of Oulu(机器视觉与信号分析中心,奥卢大学) Indian Institute of Technology Roorkee(印度理工学院罗尔基分校) State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05078 2025-05-26 cs.LG cs.AI cs.IT math.IT 79%

Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data

David Heurtel-Depeiges, Anian Ruoss, Joel Veness, Tim Genewein

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22076 2025-05-20 cs.SD cs.HC eess.AS 79%

USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis

Luca Jiang-Tao Yu, Running Zhao, Sijie Ji, Edith C. H. Ngai, Chenshu Wu

机构 * The University of Hong Kong(香港大学)

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 eess.AS

Comments Accepted by Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (ACM IMWUT/UbiComp 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏