arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2510.07592 2025-10-10 eess.AS 57%

SALAD-VAE: Semantic Audio Compression with Language-Audio Distillation

Sebastian Braun, Hannes Gamper, Dimitra Emmanouilidou

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05191 2025-10-10 cs.SD cs.AI 57%

Provable Speech Attributes Conversion via Latent Independence

Jonathan Svirsky, Ofir Lindenbaum, Uri Shaham

机构 * Bar Ilan University(巴伊兰大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07299 2025-10-09 eess.AS cs.SD 57%

Comparison of Speech Tasks in Human Expert and Machine Detection of Parkinson's Disease

Peter Plantinga, Roozbeh Sattari, Karine Marcotte, Carla Di Gironimo, Madeleine Sharp, Liziane Bouvier, Maiya Geddes, Ingrid Verduyckt, Étienne de Villers-Sidani, Mirco Ravanelli, Denise Klein

机构 * McGill University(麦吉尔大学) CRBLM Mila Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) Nouvelle Voix(新声音) Montreal Neurological Institute(蒙特利尔神经科学研究所) Concordia University(Concordia大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted to SMASH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04750 2025-10-07 cs.CL cs.SE 57%

A Low-Resource Speech-Driven NLP Pipeline for Sinhala Dyslexia Assistance

Peshala Perera, Deshan Sumanathilaka

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 11 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03336 2025-10-07 cs.SD cs.AI cs.LG 57%

Linguistic and Audio Embedding-Based Machine Learning for Alzheimer's Dementia and Mild Cognitive Impairment Detection: Insights from the PROCESS Challenge

Adharsha Sam Edwin Sam Devahi, Sohail Singh Sangha, Prachee Priyadarshinee, Jithin Thilakan, Ivan Fu Xing Tan, Christopher Johann Clarke, Sou Ka Lon, Balamurali B T, Yow Wei Quin, Chen Jer-Ming

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Hochschule für Musik Detmold(音乐学院Detmold)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02313 2025-10-03 cs.CV 57%

Clink! Chop! Thud! -- Learning Object Sounds from Real-World Interactions

Mengyu Yang, Yiming Chen, Haozheng Pei, Siddhant Agarwal, Arun Balajee Vasudevan, James Hays

机构 * Georgia Institute of Technology(佐治亚理工学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025. Project page: https://clink-chop-thud.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02110 2025-10-03 cs.SD cs.LG eess.AS 57%

SoundReactor: Frame-level Online Video-to-Audio Generation

Koichi Saito, Julian Tanke, Christian Simon, Masato Ishii, Kazuki Shimada, Zachary Novack, Zhi Zhong, Akio Hayakawa, Takashi Shibuya, Yuki Mitsufuji

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11835 2025-10-03 cs.LG cs.CV 57%

How Can Time Series Analysis Benefit From Multiple Modalities? A Survey and Outlook

Haoxin Liu, Harshavardhan Kamarthi, Zhiyuan Zhao, Shangqing Xu, Shiyu Wang, Qingsong Wen, Tom Hartvigsen, Fei Wang, B. Aditya Prakash

机构 * Georgia Institute of Technology(佐治亚理工学院) Bytedance Inc.(字节跳动公司) Squirrel AI, USA(squirrel AI 美国分公司) The University of Virginia(弗吉尼亚大学) Cornell University(康奈尔大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Github Repo: https://github.com/AdityaLab/MM4TSA Updated to include papers accepted by IJCAI25, KDD25, ICML25, NeurIPS25 4 figures or tables, 19 pages, 251 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09746 2025-10-02 cs.SD eess.AS 57%

Deep Learning for Tuberculosis Screening in a High-burden Setting using Cough Analysis and Speech Foundation Models

Ning Ma, Bahman Mirheidari, Guy J. Brown, Nsala Sanjase, Minyoi M. Maimbolwa, Solomon Chifwamba, Seke Muzazu, Monde Muyoyeta, Mary Kagujje

机构 * School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院) Centre for Infectious Disease Research in Zambia(赞比亚传染病疾病研究中心)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments submitted to IEEE Journal of Biomedical and Health Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26604 2025-10-01 cs.CV 57%

Video Object Segmentation-Aware Audio Generation

Ilpo Viertola, Vladimir Iashin, Esa Rahtu

机构 * Tampere University, Tampere, Finland(塔尔库大学) University of Oxford, Oxford, UK(牛津大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Preprint version. The Version of Record is published in DAGM GCPR 2025 proceedings with Springer Lecture Notes in Computer Science (LNCS). Updated results and resources are available at the project page: https://saganet.notion.site

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25670 2025-10-01 cs.SD cs.CV 57%

LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning

Kang Yang, Yifan Liang, Fangkun Liu, Zhenping Xie, Chengshi Zheng

机构 * School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) Institute of Acoustics Chinese Academy of Science(中国科学院声学研究所)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08753 2025-09-30 cs.CL 57%

Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling

Neil Zeghidour, Eugene Kharitonov, Manu Orsini, Václav Volhejn, Gabriel de Marmiesse, Edouard Grave, Patrick Pérez, Laurent Mazaré, Alexandre Défossez

机构 * Kyutai

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22063 2025-09-29 cs.CV cs.SD 57%

High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling

Chao Huang, Susan Liang, Yapeng Tian, Anurag Kumar, Chenliang Xu

机构 * Meta Reality Labs Research(Meta现实实验室)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Accepted to IJCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22062 2025-09-29 cs.SD eess.AS 57%

Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling

Junjie Cao, Yichen Han, Ruonan Zhang, Xiaoyang Hao, Hongxiang Li, Shuaijiang Zhao, Yue Liu, Xiao-Ping Zhng

机构 * Tsinghua University(清华大学) Peking University(北京大学) AMAP Speech(AMAP语音)

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments conference paper about TTS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21574 2025-09-29 cs.CV 57%

X-Streamer: Unified Human World Modeling with Audiovisual Interaction

You Xie, Tianpei Gu, Zenan Li, Chenxu Zhang, Guoxian Song, Xiaochen Zhao, Chao Liang, Jianwen Jiang, Hongyi Xu, Linjie Luo

机构 * ByteDance(字节跳动)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Project Page at https://byteaigc.github.io/X-Streamer

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16031 2025-09-29 cs.CV 57%

GLip: A Global-Local Integrated Progressive Framework for Robust Visual Speech Recognition

Tianyue Wang, Shuang Yang, Shiguang Shan, Xilin Chen

机构 * University of Chinese Academy of Sciences(中国科学院大学) State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全国家重点实验室,计算技术研究所)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04059 2025-09-29 cs.CL 57%

Towards an AI Musician: Synthesizing Sheet Music Problems for Musical Reasoning

Zhilin Wang, Zhe Yang, Yun Luo, Yafu Li, Xiaoye Qu, Ziqian Qiao, Haoran Zhang, Runzhe Zhan, Derek F. Wong, Jizhe Zhou, Yu Cheng

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai AI Laboratory(上海人工智能实验室) Sichuan University(四川大学) Shanghai Jiao Tong University(上海交通大学) University of Macau(澳门大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19463 2025-09-29 cs.HC cs.AI 57%

"She was useful, but a bit too optimistic": Augmenting Design with Interactive Virtual Personas

Paluck Deep, Monica Bharadhidasan, A. Baki Kocaballi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments The version accepted for publication at International Journal of Human-Computer Studies

Journal ref International Journal of Human-Computer Studies (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13068 2025-09-29 cs.SD cs.LG eess.AS 57%

On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning

Tiago Tavares, Fabio Ayres, Zhepei Wang, Paris Smaragdis

机构 * INSPER Institute of Teaching and Research(INSPER教学与研究机构) Research University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校研究大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21144 2025-09-26 cs.SD cs.AI 57%

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

Sitong Cheng, Weizhen Bian, Xinsheng Wang, Ruibin Yuan, Jianyi Chen, Shunshun Yin, Yike Guo, Wei Xue

机构 * Hong Kong University of Science and Technology(香港理工大学) Soul AI Lab(Soul AI 实验室)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07282 2025-09-26 eess.AS 57%

Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild

Jing-Tong Tzeng, Bo-Hao Su, Ya-Tse Wu, Hsing-Hang Chou, Chi-Chun Lee

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Proceedings of Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17164 2025-09-23 cs.SD eess.AS 57%

STAR: Speech-to-Audio Generation via Representation Learning

Zeyu Xie, Xuenan Xu, Yixuan Li, Mengyue Wu, Yuexian Zou

机构 * Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology, Peking University, Shenzhen(广东超高清沉浸媒体技术重点实验室,北京大学深圳校区) X-LANCE Lab, Shanghai Jiao Tong University, Shanghai(X-LANCE实验室,上海交通大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14023 2025-09-18 cs.CL cs.HC 57%

Audio-Based Crowd-Sourced Evaluation of Machine Translation Quality

Sami Ul Haq, Sheila Castilho, Yvette Graham

机构 * ADAPT Centre(ADAPT中心) Dublin City University (DCU)(都柏林城市大学) Trinity College Dublin (TCD)(三一学院都柏林)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted at WMT2025 (ENNLP) for oral presented

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10432 2025-09-18 q-bio.OT cs.AI 57%

Standards in the Preparation of Biomedical Research Metadata: A Bridge2AI Perspective

Harry Caufield, Satrajit Ghosh, Sek Wong Kong, Jillian Parker, Nathan Sheffield, Bhavesh Patel, Andrew Williams, Timothy Clark, Monica C. Munoz-Torres

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07538 2025-09-11 cs.CV 57%

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

Peijin Xie, Shun Qian, Bingquan Liu, Dexin Wang, Lin Sun, Xiangzheng Zhang

机构 * IEEE

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments 5 pages, 4 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13314 2025-09-04 cs.SD eess.AS 57%

I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception

Jiawei Zhang, Tian-Hao Zhang, Jun Wang, Jiaran Gao, Xinyuan Qian, Xu-Cheng Yin

机构 * University of Science and Technology Beijing(北京科技大学) Tencent AI Lab(腾讯AI实验室) National University of Singapore(新加坡国立大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Accepted by APSIPA ASC2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17625 2025-09-01 cs.LG cs.AI 57%

Alice's Adventures in a Differentiable Wonderland -- Volume I, A Tour of the Land

Simone Scardapane

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Companion website for additional chapters: https://www.sscardapane.it/alice-book

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20796 2025-08-29 cs.SD cs.AI 57%

Speech Emotion Recognition via Entropy-Aware Score Selection

ChenYi Chua, JunKai Wong, Chengxin Chen, Xiaoxiao Miao

机构 * Singapore Institute of Technology(新加坡理工学院) Institute Of Acoustics, Chinese Academy Of Sciences(中国科学院声学研究所) Duke Kunshan University(杜克大学昆山分校)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments The paper has been accepted by APCIPA ASC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20782 2025-08-29 eess.AS 57%

A Solution of Ultra Wideband Based High-resolution and Lossless Audio Transmission

Fengyun Zhang

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments 5 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17336 2025-08-29 cs.SD cs.AI 57%

Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework

Yunsik Kim, Yoonyoung Chung

机构 * Department of Electrical Engineering(电气工程系) Department of Semiconductor Engineering(半导体工程系) Center for Semiconductor Technology Convergence(半导体技术融合中心)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Journal ref Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏