arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2511.12662 2025-11-18 cs.CV 57%

Hi-Reco: High-Fidelity Real-Time Conversational Digital Humans

Hongbin Huang, Junwei Li, Tianxin Xie, Zhuang Li, Cekai Weng, Yaodong Yang, Yue Luo, Li Liu, Jing Tang, Zhijing Shao, Zeyu Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Prometheus Vision Technology Co., Ltd.(普罗米修斯视觉科技有限公司) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Proceedings of the Computer Graphics International 2025 (CGI'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08230 2025-11-18 cs.CL 57%

VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context

Heyang Liu, Ziyang Cheng, Yuhao Wang, Hongcheng Liu, Yiqi Li, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

Comments This article will serve as an extension of the preceding work, "VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models" (arXiv:2505.15727). Therefore, we have chosen to withdraw to avoid potential duplicate publication. We will update the previously open-sourced paper of VocalBench in several weeks to include the content of VocalBench-zh

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23447 2025-11-18 cs.CV 57%

CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition

Jongseo Lee, Joohyun Chang, Dongho Lee, Jinwoo Choi

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments Our paper has been accepted to IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08132 2025-11-14 cs.AI 57%

National Institute on Aging PREPARE Challenge: Early Detection of Cognitive Impairment Using Speech -- The SpeechCARE Solution

Maryam Zolnoori, Hossein Azadmaleki, Yasaman Haghbin, Ali Zolnour, Mohammad Javad Momeni Nezhad, Sina Rashidi, Mehdi Naserian, Elyas Esmaeili, Sepehr Karimi Arpanahi

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08670 2025-11-14 cs.HC cs.AI 57%

Once Upon an AI: Six Scaffolds for Child-AI Interaction Design, Inspired by Disney

Nomisha Kurian

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04914 2025-11-13 cs.SD cs.AI 57%

MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages

Hardik B. Sailor, Aw Ai Ti, Chen Fang Yih Nancy, Chiu Ying Lay, Ding Yang, He Yingxu, Jiang Ridong, Li Jingtao, Liao Jingyi, Liu Zhuohan, Lu Yanfeng, Ma Yi, Manas Gupta, Muhammad Huzaifah Bin Md Shahrin, Nabilah Binte Md Johan, Nattadaporn Lertcheva, Pan Chunlei, Pham Minh Duc, Siti Maryam Binte Ahmad Subaidi, Siti Umairah Binte Mohammad Salleh, Sun Shuo, Tarun Kumar Vangani, Wang Qiongqiong, Won Cheng Yi Lewis, Wong Heng Meng Jeremy, Wu Jinyang, Zhang Huayun, Zhang Longyin, Zou Xunlong

机构 * MERaLiON Team Institute for Infocomm Research (I 2 R), A*STAR, Singapore(MERaLiON团队信息与通信研究所(I 2 R),A*STAR,新加坡)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments https://huggingface.co/MERaLiON/MERaLiON-SER-v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16774 2025-11-13 cs.CL 57%

IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models

Yiming Gao, Bin Wang, Chengwei Wei, Shuo Sun, AiTi Aw

机构 * Nanyang Technological University (NTU)(南洋理工大学) MiroMind(米罗Mind) Institute for Infocomm Research (I 2 R)(信息与通信研究院) A*STAR(科技研究局)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Link: https://github.com/AudioLLMs/AudioBench/tree/main/IFEval-Audio

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01176 2025-11-04 cs.GR cs.CV cs.LG cs.SD 57%

Audio Driven Real-Time Facial Animation for Social Telepresence

Jiye Lee, Chenghui Li, Linh Tran, Shih-En Wei, Jason Saragih, Alexander Richard, Hanbyul Joo, Shaojie Bai

机构 * Seoul National University(首尔国立大学) Codec Avatars Lab, Meta(Meta语义编码实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments SIGGRAPH Asia 2025. Project page: https://jiyewise.github.io/projects/AudioRTA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27272 2025-11-03 cs.HC eess.AS eess.SP 57%

Inferring trust in recommendation systems from brain, behavioural, and physiological data

Vincent K. M. Cheung, Pei-Cheng Shih, Masato Hirano, Masataka Goto, Shinichi Furuya

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14755 2025-10-30 cs.DC cs.AI 57%

Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models

Daoyuan Chen, Yilun Huang, Xuchen Pan, Nana Jiang, Haibin Wang, Yilei Zhang, Ce Ge, Yushuo Chen, Wenhao Zhang, Zhijian Ma, Jun Huang, Wei Lin, Yaliang Li, Bolin Ding, Jingren Zhou

机构 * Alibaba Group(阿里巴巴集团)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 (Spotlight). 43 pages, 16 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25199 2025-10-30 cs.CV 57%

AI-Powered Early Detection of Critical Diseases using Image Processing and Audio Analysis

Manisha More, Kavya Bhand, Kaustubh Mukdam, Kavya Sharma, Manas Kawtikwar, Hridayansh Kaware, Prajwal Kavhar

机构 * Dept. of Computer Engineering(计算机工程系) Vishwakarma Institute of Technology(维斯瓦卡arma技术学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19902 2025-10-30 cs.CL 57%

WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction

Binbin Zhang, Chengdong Liang, Shuai Wang, Xuelong Geng, Zhao Guo, Haoyu Li, Hao Yin, Xipeng Yang, Pengshen Zhang, Changwei Ma, Lei Xie

机构 * Northwestern Polytechnical University(西北工业大学) Nanjing University(南京大学) Shanghai Jiao Tong University(上海交通大学) GuaSemi Speech A Team(GuaSemi语音团队) WeNet Community(WeNet社区)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24628 2025-10-29 cs.CL 57%

"Mm, Wat?" Detecting Other-initiated Repair Requests in Dialogue

Anh Ngo, Nicolas Rollet, Catherine Pelachaud, Chloe Clavel

机构 * ALMAnaCH, INRIA Paris(ALMAnaCH,INRIA巴黎) Télécom Paris, SES, Institut Polytechnique de Paris, I3-CNRS(Télécom巴黎,SES,巴黎高等理工学院,I3-CNRS) Télécom Paris, LTCI, Institut Polytechnique de Paris(Télécom巴黎,LTCI,巴黎高等理工学院) CNRS, ISIR, Sorbonne University(CNRS,ISIR,索邦大学) ISIR, Sorbonne University(ISIR,索邦大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18195 2025-10-29 eess.AS cs.LG cs.SD 57%

Acoustic and Machine Learning Methods for Speech-Based Suicide Risk Assessment: A Systematic Review

Ambre Marie, Marine Garnier, Thomas Bertin, Laura Machart, Guillaume Dardenne, Gwenolé Quellec, Sofian Berrouiguet

机构 * University of Western Brittany(西部布列塔尼大学) Psychiatry Department, CHU Brest(布列塔尼大学精神科部门)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Preprint version of a manuscript submitted to the Journal of Affective Disorders

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21581 2025-10-27 cs.CV cs.SD 57%

Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video

Ciara Rowles, Varun Jampani, Simon Donné, Shimon Vainer, Julian Parker, Zach Evans

机构 * Stability AI

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments Project Page: https://stability-ai.github.io/foleycontrol.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21043 2025-10-27 cs.CL cs.RO 57%

Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction

Sam O'Connor Russell, Naomi Harte

机构 * ADAPT Centre, School of Engineering, Trinity College Dublin(ADAPT中心、工程学院、都柏林信任学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted to ACL 2025, Findings of the Association for Computational Linguistics

Journal ref In Findings of the Association for Computational Linguistics: ACL 2025, pages 209--221, Vienna, Austria. Association for Computational Linguistics, 10.18653/v1/2025.findings-acl.12

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18972 2025-10-27 cs.LG cs.AI 57%

Deep Insights into Cognitive Decline: A Survey of Leveraging Non-Intrusive Modalities with Deep Learning Techniques

David Ortiz-Perez, Manuel Benavent-Lledo, Jose Garcia-Rodriguez, David Tomás, M. Flores Vizcaya-Moreno

机构 * Dept. of Computer Science and Technology(计算机科学与技术系) University of Alicante(阿尔瓦登特大学) Unit of Clinical Nursing Research(临床护理研究单位) Faculty of Health Sciences(健康科学学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Journal ref Applied Soft Computing, Vol. 184, 2025, Article 113787

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20867 2025-10-27 cs.LG cs.AI 57%

Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards

Jiajun Fan, Roger Ren, Jingyuan Li, Rahul Pandey, Prashanth Gurunath Shivakumar, Ivan Bulyko, Ankur Gandhe, Ge Liu, Yile Gu

机构 * Amazon(亚马逊) Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(计算机与数据科学学院,伊利诺伊大学厄巴纳-香槟分校)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 49 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19144 2025-10-23 cs.CL 57%

Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges

Cheng Huang, Nyima Tashi, Fan Gao, Yutong Liu, Jiahao Li, Hao Tian, Siyang Jiang, Thupten Tsering, Ban Ma-bao, Renzeg Duojie, Gadeng Luosang, Rinchen Dongrub, Dorje Tashi, Jin Zhang, Xiao Feng, Hao Wang, Jie Tang, Guojie Tang, Xiangxiang Wang, Jia Zhang, Tsengdar Lee, Yongbin Yu

机构 * University of Electronic Science and Technology of China(电子科技大学) Southern Methodist University(南方 Methodist 大学) The City University of Hong Kong(香港城市大学) The Hong Kong Polytechnic University(香港理工大学) The Chinese University of Hong Kong(香港中文大学) University of Connecticut(康涅狄格大学) Tsinghua University(清华大学) University of Texas at Arlington(德克萨斯大学阿灵顿分校) University of Chinese Academy of Sciences(中国科学院大学) Tibet University(西藏大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17617 2025-10-21 cs.HC cs.CV 57%

ImaGGen: Zero-Shot Generation of Co-Speech Semantic Gestures Grounded in Language and Image Input

Hendric Voss, Stefan Kopp

机构 * Social Cognitive Systems Group, Bielefeld University(比勒菲尔德大学社会认知系统小组)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15757 2025-10-20 cs.LG cs.CV cs.NE 57%

Poultry Farm Intelligence: An Integrated Multi-Sensor AI Platform for Enhanced Welfare and Productivity

Pieris Panagi, Savvas Karatsiolis, Kyriacos Mosphilis, Nicholas Hadjisavvas, Andreas Kamilaris, Nicolas Nicolaou, Efstathios Stavrakis, Vassilis Vassiliades

机构 * CYENS - Centre of Excellence(CYENS 卓越中心) Algolysis Ltd(Algolysis 公司) Department of Computer Science(计算机科学系)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10248 2025-10-20 cs.LG cs.AI 57%

Reasoning-Enhanced Large Language Models for Molecular Property Prediction

Jiaxi Zhuang, Yaorui Shi, Jue Hou, Yunong He, Mingwei Ye, Mingjun Xu, Yuming Su, Linfeng Zhang, Ying Qian, Linfeng Zhang, Guolin Ke, Hengxing Cai

机构 * DP Technology(DP技术公司) Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02569 2025-10-20 cs.CL 57%

Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models

Tolúlopé Ògúnrèmí, Christopher D. Manning, Dan Jurafsky, Karen Livescu

机构 * Stanford University(斯坦福大学) Toyota Technological Institute at Chicago(芝加哥丰田技术研究所)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00059 2025-10-20 cs.CV cs.LG 57%

Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models

Rui Hu, Delai Qiu, Shuyu Wei, Jiaming Zhang, Yining Wang, Shengping Liu, Jitao Sang

机构 * Beijing Key Lab of Traffic Data Analysis and Mining(北京交通数据挖掘重点实验室) Beijing Jiaotong University(北京交通大学) Unisound AI Technology Co., Ltd.(Unisound人工智能技术有限公司) Peng Cheng Lab(鹏城实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted to ACL 2025 Findings

Journal ref Findings of the Association for Computational Linguistics: ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13344 2025-10-16 cs.SD cs.CL 57%

UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE

Zhenyu Liu, Yunxin Li, Xuanyu Zhang, Qixun Teng, Shenyuan Jiang, Xinyu Chen, Haoyuan Shi, Jinchao Li, Qi Wang, Haolan Chen, Fanbo Meng, Mingjun Zhao, Yu Xu, Yancheng He, Baotian Hu, Min Zhang

机构 * Department of Computer Science and Technology, Harbin Institute of Technology(计算机科学与技术系,哈尔滨工业大学) Shenzhen Loop Area Institute(深圳河套学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05803 2025-10-16 cs.GR cs.CV 57%

PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis

Yihuan Huang, Jiajun Liu, Yanzhen Ren, Jun Xue, Wuyang Liu, Zongkun Sun

机构 * Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航空航天信息安全部分和可信计算重点实验室、教育部、网络安全科学与工程学院、武汉大学) School of Cyber Science and Engineering, Wuhan University(网络安全科学与工程学院、武汉大学) School of Police Information, Shandong Police College(警务信息学院、山东警察学院)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11454 2025-10-14 cs.SD cs.AI 57%

Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning

Kuan-Yi Lee, Tsung-En Lin, Hung-Yi Lee

机构 * National Taiwan University(国立台湾大学) ASUS Open Cloud Infrastructure Software Center(ASUS开放云基础设施软件中心)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments 9pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10329 2025-10-14 cs.CL 57%

End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs

Nam Luu, Ondřej Bojar

机构 * Charles University(查尔斯大学) Faculty of Mathematics and Physics(数学与物理系) Institute of Formal and Applied Linguistics(形式与应用语言学研究所)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10173 2025-10-14 cs.HC cs.CY cs.SD eess.AS 57%

Chord Colourizer: A Near Real-Time System for Visualizing Musical Key

Paul Haimes

机构 * Ritsumeikan University(立命馆大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Author copy. This paper is in press for presentation at ADADA 2025. Please cite as: Haimes, P. (in press). Chord Colourizer: A near real-time system for visualizing musical key. In Proceedings of the 23rd International Conference of Asia Digital Art and Design Association (ADADA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07881 2025-10-10 cs.CL 57%

CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching

Heyang Liu, Yuhao Wang, Ziyang Cheng, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Ant Group(蚂蚁集团)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏