arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2510.27272 2025-11-03 cs.HC eess.AS eess.SP 57%

Inferring trust in recommendation systems from brain, behavioural, and physiological data

Vincent K. M. Cheung, Pei-Cheng Shih, Masato Hirano, Masataka Goto, Shinichi Furuya

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14755 2025-10-30 cs.DC cs.AI 57%

Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models

Daoyuan Chen, Yilun Huang, Xuchen Pan, Nana Jiang, Haibin Wang, Yilei Zhang, Ce Ge, Yushuo Chen, Wenhao Zhang, Zhijian Ma, Jun Huang, Wei Lin, Yaliang Li, Bolin Ding, Jingren Zhou

机构 * Alibaba Group(阿里巴巴集团)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 (Spotlight). 43 pages, 16 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25199 2025-10-30 cs.CV 57%

AI-Powered Early Detection of Critical Diseases using Image Processing and Audio Analysis

Manisha More, Kavya Bhand, Kaustubh Mukdam, Kavya Sharma, Manas Kawtikwar, Hridayansh Kaware, Prajwal Kavhar

机构 * Dept. of Computer Engineering(计算机工程系) Vishwakarma Institute of Technology(维斯瓦卡arma技术学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19902 2025-10-30 cs.CL 57%

WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction

Binbin Zhang, Chengdong Liang, Shuai Wang, Xuelong Geng, Zhao Guo, Haoyu Li, Hao Yin, Xipeng Yang, Pengshen Zhang, Changwei Ma, Lei Xie

机构 * Northwestern Polytechnical University(西北工业大学) Nanjing University(南京大学) Shanghai Jiao Tong University(上海交通大学) GuaSemi Speech A Team(GuaSemi语音团队) WeNet Community(WeNet社区)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24628 2025-10-29 cs.CL 57%

"Mm, Wat?" Detecting Other-initiated Repair Requests in Dialogue

Anh Ngo, Nicolas Rollet, Catherine Pelachaud, Chloe Clavel

机构 * ALMAnaCH, INRIA Paris(ALMAnaCH,INRIA巴黎) Télécom Paris, SES, Institut Polytechnique de Paris, I3-CNRS(Télécom巴黎,SES,巴黎高等理工学院,I3-CNRS) Télécom Paris, LTCI, Institut Polytechnique de Paris(Télécom巴黎,LTCI,巴黎高等理工学院) CNRS, ISIR, Sorbonne University(CNRS,ISIR,索邦大学) ISIR, Sorbonne University(ISIR,索邦大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18195 2025-10-29 eess.AS cs.LG cs.SD 57%

Acoustic and Machine Learning Methods for Speech-Based Suicide Risk Assessment: A Systematic Review

Ambre Marie, Marine Garnier, Thomas Bertin, Laura Machart, Guillaume Dardenne, Gwenolé Quellec, Sofian Berrouiguet

机构 * University of Western Brittany(西部布列塔尼大学) Psychiatry Department, CHU Brest(布列塔尼大学精神科部门)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Preprint version of a manuscript submitted to the Journal of Affective Disorders

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21581 2025-10-27 cs.CV cs.SD 57%

Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video

Ciara Rowles, Varun Jampani, Simon Donné, Shimon Vainer, Julian Parker, Zach Evans

机构 * Stability AI

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Project Page: https://stability-ai.github.io/foleycontrol.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21043 2025-10-27 cs.CL cs.RO 57%

Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction

Sam O'Connor Russell, Naomi Harte

机构 * ADAPT Centre, School of Engineering, Trinity College Dublin(ADAPT中心、工程学院、都柏林信任学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted to ACL 2025, Findings of the Association for Computational Linguistics

Journal ref In Findings of the Association for Computational Linguistics: ACL 2025, pages 209--221, Vienna, Austria. Association for Computational Linguistics, 10.18653/v1/2025.findings-acl.12

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18972 2025-10-27 cs.LG cs.AI 57%

Deep Insights into Cognitive Decline: A Survey of Leveraging Non-Intrusive Modalities with Deep Learning Techniques

David Ortiz-Perez, Manuel Benavent-Lledo, Jose Garcia-Rodriguez, David Tomás, M. Flores Vizcaya-Moreno

机构 * Dept. of Computer Science and Technology(计算机科学与技术系) University of Alicante(阿尔瓦登特大学) Unit of Clinical Nursing Research(临床护理研究单位) Faculty of Health Sciences(健康科学学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Journal ref Applied Soft Computing, Vol. 184, 2025, Article 113787

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20867 2025-10-27 cs.LG cs.AI 57%

Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards

Jiajun Fan, Roger Ren, Jingyuan Li, Rahul Pandey, Prashanth Gurunath Shivakumar, Ivan Bulyko, Ankur Gandhe, Ge Liu, Yile Gu

机构 * Amazon(亚马逊) Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(计算机与数据科学学院,伊利诺伊大学厄巴纳-香槟分校)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 49 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19144 2025-10-23 cs.CL 57%

Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges

Cheng Huang, Nyima Tashi, Fan Gao, Yutong Liu, Jiahao Li, Hao Tian, Siyang Jiang, Thupten Tsering, Ban Ma-bao, Renzeg Duojie, Gadeng Luosang, Rinchen Dongrub, Dorje Tashi, Jin Zhang, Xiao Feng, Hao Wang, Jie Tang, Guojie Tang, Xiangxiang Wang, Jia Zhang, Tsengdar Lee, Yongbin Yu

机构 * University of Electronic Science and Technology of China(电子科技大学) Southern Methodist University(南方 Methodist 大学) The City University of Hong Kong(香港城市大学) The Hong Kong Polytechnic University(香港理工大学) The Chinese University of Hong Kong(香港中文大学) University of Connecticut(康涅狄格大学) Tsinghua University(清华大学) University of Texas at Arlington(德克萨斯大学阿灵顿分校) University of Chinese Academy of Sciences(中国科学院大学) Tibet University(西藏大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17617 2025-10-21 cs.HC cs.CV 57%

ImaGGen: Zero-Shot Generation of Co-Speech Semantic Gestures Grounded in Language and Image Input

Hendric Voss, Stefan Kopp

机构 * Social Cognitive Systems Group, Bielefeld University(比勒菲尔德大学社会认知系统小组)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15757 2025-10-20 cs.LG cs.CV cs.NE 57%

Poultry Farm Intelligence: An Integrated Multi-Sensor AI Platform for Enhanced Welfare and Productivity

Pieris Panagi, Savvas Karatsiolis, Kyriacos Mosphilis, Nicholas Hadjisavvas, Andreas Kamilaris, Nicolas Nicolaou, Efstathios Stavrakis, Vassilis Vassiliades

机构 * CYENS - Centre of Excellence(CYENS 卓越中心) Algolysis Ltd(Algolysis 公司) Department of Computer Science(计算机科学系)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10248 2025-10-20 cs.LG cs.AI 57%

Reasoning-Enhanced Large Language Models for Molecular Property Prediction

Jiaxi Zhuang, Yaorui Shi, Jue Hou, Yunong He, Mingwei Ye, Mingjun Xu, Yuming Su, Linfeng Zhang, Ying Qian, Linfeng Zhang, Guolin Ke, Hengxing Cai

机构 * DP Technology(DP技术公司) Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02569 2025-10-20 cs.CL 57%

Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models

Tolúlopé Ògúnrèmí, Christopher D. Manning, Dan Jurafsky, Karen Livescu

机构 * Stanford University(斯坦福大学) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00059 2025-10-20 cs.CV cs.LG 57%

Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models

Rui Hu, Delai Qiu, Shuyu Wei, Jiaming Zhang, Yining Wang, Shengping Liu, Jitao Sang

机构 * Beijing Key Lab of Traffic Data Analysis and Mining(北京交通数据挖掘重点实验室) Beijing Jiaotong University(北京交通大学) Unisound AI Technology Co., Ltd.(Unisound人工智能技术有限公司) Peng Cheng Lab(鹏城实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to ACL 2025 Findings

Journal ref Findings of the Association for Computational Linguistics: ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13344 2025-10-16 cs.SD cs.CL 57%

UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE

Zhenyu Liu, Yunxin Li, Xuanyu Zhang, Qixun Teng, Shenyuan Jiang, Xinyu Chen, Haoyuan Shi, Jinchao Li, Qi Wang, Haolan Chen, Fanbo Meng, Mingjun Zhao, Yu Xu, Yancheng He, Baotian Hu, Min Zhang

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05803 2025-10-16 cs.GR cs.CV 57%

PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis

Yihuan Huang, Jiajun Liu, Yanzhen Ren, Jun Xue, Wuyang Liu, Zongkun Sun

机构 * Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航空航天信息安全部分和可信计算重点实验室、教育部、网络安全科学与工程学院、武汉大学) School of Cyber Science and Engineering, Wuhan University(网络安全科学与工程学院、武汉大学) School of Police Information, Shandong Police College(警务信息学院、山东警察学院)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11454 2025-10-14 cs.SD cs.AI 57%

Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning

Kuan-Yi Lee, Tsung-En Lin, Hung-Yi Lee

机构 * National Taiwan University(国立台湾大学) ASUS Open Cloud Infrastructure Software Center(ASUS开放云基础设施软件中心)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 9pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10329 2025-10-14 cs.CL 57%

End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs

Nam Luu, Ondřej Bojar

机构 * Charles University(查尔斯大学) Faculty of Mathematics and Physics(数学与物理系) Institute of Formal and Applied Linguistics(形式与应用语言学研究所)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10173 2025-10-14 cs.HC cs.CY cs.SD eess.AS 57%

Chord Colourizer: A Near Real-Time System for Visualizing Musical Key

Paul Haimes

机构 * Ritsumeikan University(立命馆大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Author copy. This paper is in press for presentation at ADADA 2025. Please cite as: Haimes, P. (in press). Chord Colourizer: A near real-time system for visualizing musical key. In Proceedings of the 23rd International Conference of Asia Digital Art and Design Association (ADADA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07881 2025-10-10 cs.CL 57%

CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching

Heyang Liu, Yuhao Wang, Ziyang Cheng, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Ant Group(蚂蚁集团)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07592 2025-10-10 eess.AS 57%

SALAD-VAE: Semantic Audio Compression with Language-Audio Distillation

Sebastian Braun, Hannes Gamper, Dimitra Emmanouilidou

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05191 2025-10-10 cs.SD cs.AI 57%

Provable Speech Attributes Conversion via Latent Independence

Jonathan Svirsky, Ofir Lindenbaum, Uri Shaham

机构 * Bar Ilan University(巴伊兰大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07299 2025-10-09 eess.AS cs.SD 57%

Comparison of Speech Tasks in Human Expert and Machine Detection of Parkinson's Disease

Peter Plantinga, Roozbeh Sattari, Karine Marcotte, Carla Di Gironimo, Madeleine Sharp, Liziane Bouvier, Maiya Geddes, Ingrid Verduyckt, Étienne de Villers-Sidani, Mirco Ravanelli, Denise Klein

机构 * McGill University(麦吉尔大学) CRBLM Mila Quebec AI Institute(魁北克人工智能研究所) Université de Montréal(蒙特利尔大学) Nouvelle Voix(新声音) Montreal Neurological Institute(蒙特利尔神经科学研究所) Concordia University(Concordia大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted to SMASH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04750 2025-10-07 cs.CL cs.SE 57%

A Low-Resource Speech-Driven NLP Pipeline for Sinhala Dyslexia Assistance

Peshala Perera, Deshan Sumanathilaka

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 11 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03336 2025-10-07 cs.SD cs.AI cs.LG 57%

Linguistic and Audio Embedding-Based Machine Learning for Alzheimer's Dementia and Mild Cognitive Impairment Detection: Insights from the PROCESS Challenge

Adharsha Sam Edwin Sam Devahi, Sohail Singh Sangha, Prachee Priyadarshinee, Jithin Thilakan, Ivan Fu Xing Tan, Christopher Johann Clarke, Sou Ka Lon, Balamurali B T, Yow Wei Quin, Chen Jer-Ming

机构 * Singapore University of Technology and Design(新加坡科技设计大学) Hochschule für Musik Detmold(音乐学院Detmold)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02313 2025-10-03 cs.CV 57%

Clink! Chop! Thud! -- Learning Object Sounds from Real-World Interactions

Mengyu Yang, Yiming Chen, Haozheng Pei, Siddhant Agarwal, Arun Balajee Vasudevan, James Hays

机构 * Georgia Institute of Technology(佐治亚理工学院) Carnegie Mellon University(卡内基梅隆大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments ICCV 2025. Project page: https://clink-chop-thud.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02110 2025-10-03 cs.SD cs.LG eess.AS 57%

SoundReactor: Frame-level Online Video-to-Audio Generation

Koichi Saito, Julian Tanke, Christian Simon, Masato Ishii, Kazuki Shimada, Zachary Novack, Zhi Zhong, Akio Hayakawa, Takashi Shibuya, Yuki Mitsufuji

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11835 2025-10-03 cs.LG cs.CV 57%

How Can Time Series Analysis Benefit From Multiple Modalities? A Survey and Outlook

Haoxin Liu, Harshavardhan Kamarthi, Zhiyuan Zhao, Shangqing Xu, Shiyu Wang, Qingsong Wen, Tom Hartvigsen, Fei Wang, B. Aditya Prakash

机构 * Georgia Institute of Technology(佐治亚理工学院) Bytedance Inc.(字节跳动公司) Squirrel AI, USA(squirrel AI 美国分公司) The University of Virginia(弗吉尼亚大学) Cornell University(康奈尔大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Github Repo: https://github.com/AdityaLab/MM4TSA Updated to include papers accepted by IJCAI25, KDD25, ICML25, NeurIPS25 4 figures or tables, 19 pages, 251 references

详情

展开后加载摘要…

URL PDF HTML 收藏