arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2508.06382 2025-08-11 cs.CV 57%

Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning

Xiangyu Wu, Feng Yu, Yang Yang, Jianfeng Lu

机构 * Nanjing University of Science

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted for publication at ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05978 2025-08-11 cs.SD cs.AI cs.LG 57%

DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching

Wei Chen, Binzhu Sha, Dan Luo, Jing Yang, Zhuo Wang, Fan Fan, Zhiyong Wu

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 音频语音多模态 :any-to-any(abstract);分类 cs.AI

Comments Accepted by INTERSPEECH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04585 2025-08-08 eess.AS 57%

UniTalker: Conversational Speech-Visual Synthesis

Yifan Hu, Rui Liu, Yi Ren, Xiang Yin, Haizhou Li

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 15 pages, 8 figures, Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03536 2025-08-08 eess.AS 57%

Overview of Automatic Speech Analysis and Technologies for Neurodegenerative Disorders: Diagnosis and Assistive Applications

Shakeel A. Sheikh, Md. Sahidullah, Ina Kodrasi

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Published in IEEE Journal of Selected Topics in Signal Processing

Journal ref https://ieeexplore.ieee.org/abstract/document/11086511/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04481 2025-08-07 cs.LG cs.HC cs.NE cs.SD eess.AS 57%

Emotion Detection Using Conditional Generative Adversarial Networks (cGAN): A Deep Learning Approach

Anushka Srivastava

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 3 pages, 2 tables, submitted for arXiv preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00205 2025-08-04 cs.CV 57%

Learning Personalised Human Internal Cognition from External Expressive Behaviours for Real Personality Recognition

Xiangyu Kong, Hengde Zhu, Haoqin Sun, Zhihao Guo, Jiayan Gu, Xinyi Ni, Wei Zhang, Shizhe Liu, Siyang Song

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23590 2025-08-01 cs.SD eess.AS 57%

Identifying Hearing Difficulty Moments in Conversational Audio

Jack Collins, Adrian Buzea, Chris Collier, Alejandro Ballesta Rosen, Julian Maclaren, Richard F. Lyon, Simon Carlile

机构 * Google Research Australia(谷歌澳大利亚研究)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20999 2025-07-31 cs.SD cs.GR eess.AS 57%

Text-Driven Voice Conversion via Latent State-Space Modeling

Wen Li, Sofia Martinez, Priyanka Shah

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11308 2025-07-31 eess.AS eess.SP 57%

Uncovering the role of semantic and acoustic cues in normal and dichotic listening

Sai Samrat Kankanala, Akshara Soman, Sriram Ganapathy

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments 10 Pages, 6 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22008 2025-07-30 cs.CV 57%

VeS: Teaching Pixels to Listen Without Supervision

Sajay Raj

机构 * Indian Institute of Technology, Madras(印度理工学院马德拉斯分校)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV

Comments 6 pages, 1 figure, 1 table. Code and models are released

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18380 2025-07-28 cs.AI cs.LG 57%

RedactOR: An LLM-Powered Framework for Automatic Clinical Data De-Identification

Praphul Singh, Charlotte Dzialo, Jangwon Kim, Sumana Srivatsa, Irfan Bulu, Sri Gadde, Krishnaram Kenthapadi

机构 * Oracle Health & AI(Oracle健康与人工智能)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI

Comments Accepted to ACL 2025 Industry Track. To appear

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07926 2025-07-25 cs.RO cs.CV cs.LG 57%

Learning Gentle Grasping Using Vision, Sound, and Touch

Ken Nakahara, Roberto Calandra

机构 * Learning, Adaptive Systems, and Robotics (LASR) Lab, TU Dresden(图腾德斯登技术大学学习、自适应系统与机器人实验室) Center for Tactile Internet with Human-in-the-Loop (CeTI)(人机协同触觉互联网中心)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV

Comments 8 pages. Accepted by 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15007 2025-07-23 cs.PL cs.CL 57%

Hear Your Code Fail, Voice-Assisted Debugging for Python

Sayed Mahbub Hasan Amiri, Md. Mainul Islam, Mohammad Shakhawat Hossen, Sayed Majhab Hasan Amiri, Mohammad Shawkat Ali Mamun, Sk. Humaun Kabir, Naznin Akter

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments 35 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.10266 2025-07-21 cs.CL cs.LG 57%

psifx -- Psychological and Social Interactions Feature Extraction Package

Guillaume Rochette, Mathieu Rochat, Matthew J. Vowels

机构 * UNIL ETH Zurich(苏黎世联邦理工学院) Institute of Neuroinformatics(神经信息学研究所) Kivira Health(Kivira健康)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19598 2025-07-11 cs.CL 57%

Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical Study

Guanyu Hou, Jiaming He, Yinhang Zhou, Ji Guo, Yitong Qiao, Rui Zhang, Wenbo Jiang

机构 * University of Electronic Science and Technology of China(电子科技大学) Chengdu University of Technology(成都理工大学) Sun Yat-Sen University(中山大学)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00718 2025-07-11 cs.LG cs.SD eess.AS 57%

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models

Isha Gupta, David Khachaturov, Robert Mullins

机构 * ETH Zürich(苏黎世联邦理工学院) University of Cambridge(剑桥大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22116 2025-06-30 cs.RO cs.CV 57%

Evaluating Pointing Gestures for Target Selection in Human-Robot Collaboration

Noora Sassali, Roel Pieters

机构 * Cognitive Robotics group, Unit of Automation Technology and Mechanical Engineering, Tampere University(认知机器人组、自动化技术与机械工程单位、塔尔库大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

Comments Accepted by the 2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19494 2025-06-26 cs.CL cs.LG 57%

Graph Linearization Methods for Reasoning on Graphs with Large Language Models

Christos Xypolopoulos, Guokan Shang, Xiao Fei, Giannis Nikolentzos, Hadi Abdine, Iakovos Evdaimon, Michail Chatzianastasis, Giorgos Stamou, Michalis Vazirgiannis

机构 * Ecole Polytechnique(巴黎高等理工学院) MBZUAI(马克斯·普朗克人工智能研究所) NTUA(希腊国家技术研究中心) University of Peloponnese(希腊皮洛斯大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19398 2025-06-25 cs.SD eess.AS 57%

ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Shengkui Zhao, Zexu Pan, Bin Ma

机构 * Tongyi Lab(通义实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments accepted by Interspeech 2025, 5 pages, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17815 2025-06-24 cs.SD eess.AS 57%

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding

Julien Guinot, Alain Riou, Elio Quinton, György Fazekas

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted to ISMIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24446 2025-06-24 cs.SD eess.AS 57%

Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge

Longjie Luo, Shenghui Lu, Lin Li, Qingyang Hong

机构 * School of Electronic Science and Engineering(电子科学与工程学院) School of Informatics(信息学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted by InterSpeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16833 2025-06-23 cs.SD eess.AS 57%

Hybrid-Sep: Language-queried audio source separation via pre-trained Model Fusion and Adversarial Diffusion Training

Jianyuan Feng, Guangzheng Li, Yangfei Xu

专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS

Comments Submitted to WASAA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16020 2025-06-23 cs.SD eess.AS 57%

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schrödinger Bridge

Zijing Zhao, Kai Wang, Hao Huang, Ying Hu, Liang He, Jichen Yang

机构 * School of Computer Science(计算机科学学院) Department of Electronic Engineering(电子工程系) School of Cyber Security(网络安全学院)

专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS

Comments Accepted by Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14495 2025-06-18 cs.CV 57%

I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs

Yu Qi, Lipeng Gu, Honghua Chen, Liangliang Nan, Mingqiang Wei

机构 * School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院,南京航空航天大学) Urban Data Science Section, Delft University of Technology(都市数据科学部门,代尔夫特理工大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13990 2025-06-18 cs.AI 57%

Machine Mirages: Defining the Undefined

Hamidou Tembine

机构 * Department of Electrical Engineering and Computer Science, School of Engineering, UQTR, Quebec, Canada(电气工程与计算机科学系,工程学院,UQTR,魁北克,加拿大)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

Comments Submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19510 2025-06-17 cs.CL 57%

Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning

Yexing Du, Youcheng Pan, Ziyang Ma, Bo Yang, Yifan Yang, Keqi Deng, Xie Chen, Yang Xiang, Ming Liu, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Pengcheng Laboratory(鹏城实验室) Shanghai Jiao Tong University(上海交通大学) University of Cambridge(剑桥大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments Accepted in ACL 2025 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11344 2025-06-16 cs.CL 57%

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models

Peilin Wu, Jinho D. Choi

机构 * Emory University(埃默里大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09549 2025-06-12 eess.AS cs.SD eess.SP 57%

A Study on Speech Assessment with Visual Cues

Shafique Ahmed, Ryandhimas E. Zezario, Nasir Saleem, Amir Hussain, Hsin-Min Wang, Yu Tsao

机构 * Research Center for Information Technology Innovation(信息技术创新研究中心) Institute of Information Science(信息科学研究院)

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

Comments Accepted to Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05335 2025-06-10 cs.SD eess.AS 57%

FLAM: Frame-Wise Language-Audio Modeling

Yusong Wu, Christos Tsirigotis, Ke Chen, Cheng-Zhi Anna Huang, Aaron Courville, Oriol Nieto, Prem Seetharaman, Justin Salamon

专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS

Comments Accepted at ICML 2025 V2: fixed small typo on eq. 15 and eq. 17

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12408 2025-06-06 cs.CL 57%

On the Robust Approximation of ASR Metrics

Abdul Waheed, Hanin Atwany, Rita Singh, Bhiksha Raj

机构 * Carnegie Mellon University(卡内基梅隆大学) MBZUAI(穆斯林人工智能研究所)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL

Comments ACL 2025 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏