arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46430 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2405.07354 2025-08-26 cs.SD cs.IR cs.LG cs.MM eess.AS 62%

SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset

Sushant Gautam, Mehdi Houshmand Sarkhoosh, Jan Held, Cise Midoglu, Anthony Cioppa, Silvio Giancola, Vajira Thambawita, Michael A. Riegler, Pål Halvorsen, Mubarak Shah

机构 * SimulaMet OsloMet Forzasys University of Central Florida(佛罗里达中央大学) University of Liège(列日大学) KAUST(科威特大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16835 2025-08-25 eess.AS cs.CL 62%

Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems

Rumi Allbert, Nima Yazdani, Ali Ansari, Aruj Mahajan, Amirhossein Afsharrad, Seyed Shahabeddin Mousavi

机构 * University of Southern California(南加州大学) Stanford University(斯坦福大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24115 2025-08-19 cs.CL cs.MM 62%

TeleAntiFraud-28k: An Audio-Text Slow-Thinking Dataset for Telecom Fraud Detection

Zhiming Ma, Peidong Wang, Minhua Huang, Jingpeng Wang, Kai Wu, Xiangzhao Lv, Yachun Pang, Yin Yang, Wenjie Tang, Yuchen Kang

机构 * China Mobile Internet Company Ltd.(中国移动互联网有限公司) Northeastern University(东北大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11187 2025-08-18 eess.AS cs.CL cs.SD 62%

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style

Wonjune Kang, Deb Roy

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS

Comments Accepted to ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06117 2025-08-15 cs.CL cs.AI 62%

Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding

Fabian David Schmidt, Ivan Vulić, Goran Glavaš, David Ifeoluwa Adelani

机构 * Center For Artificial Intelligence and Data Science, University of Würzburg(人工智能与数据科学中心,乌尔姆大学) Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学) Mila, McGill University and Canada CIFAR AI Chair(Mila,麦吉尔大学及加拿大CIFAR人工智能主席)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07973 2025-08-12 cs.SD cs.CL eess.AS 62%

Joint Transcription of Acoustic Guitar Strumming Directions and Chords

Sebastian Murgul, Johannes Schimper, Michael Heizmann

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted to the 26th International Society for Music Information Retrieval Conference (ISMIR), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21772 2025-08-11 cs.MM cs.AI 62%

Solving Copyright Infringement on Short Video Platforms: Novel Datasets and an Audio Restoration Deep Learning Pipeline

Minwoo Oh, Minsu Park, Eunil Park

机构 * Department of MetaBioHealth, Sungkyunkwan University, Korea(韩国成均馆大学代谢生物健康系) Department of Applied Artificial Intelligence, Sungkyunkwan University, Korea(韩国成均馆大学应用人工智能系) Department of Computer Science and Engineering, Jaume I University, Spain(西班牙伊萨贝拉大学计算机科学与工程系)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、cs.MM

Comments Accepted for publication at IJCAI 2025. 9 pages, 4 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22053 2025-08-06 cs.SD cs.MA cs.MM eess.AS 62%

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation

Yan Rong, Jinting Wang, Guangzhi Lei, Shan Yang, Li Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Tencent AI Lab(腾讯人工智能实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01915 2025-08-05 cs.CV cs.ET cs.HC cs.LG cs.SD eess.AS 62%

EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses

Akshay Paruchuri, Sinan Hersek, Lavisha Aggarwal, Qiao Yang, Xin Liu, Achin Kulshrestha, Andrea Colaco, Henry Fuchs, Ishan Chatterjee

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Google(谷歌)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、eess.AS

Comments 15 pages, 6 figres, 6 tables. Accepted to ISMAR 2025 as a TVCG journal paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06690 2025-08-05 cs.SD cs.AI cs.MM 62%

Benchmarking Sub-Genre Classification For Mainstage Dance Music

Hongzhi Shu, Xinglin Li, Hongyu Jiang, Minghao Fu, Xinyu Li

机构 * Johns Hopkins University(约翰霍普金斯大学) Southeast University(东南大学) National University of Defense Technology(国防科技大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、cs.MM

Comments WASPAA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00391 2025-08-04 cs.CV eess.AS 62%

Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition

Guanjie Huang, Danny H. K. Tsang, Shan Yang, Guangzhi Lei, Li Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tencent AI Lab(腾讯AI实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、eess.AS

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00160 2025-08-04 cs.HC cs.AI cs.SD eess.AS 62%

DeformTune: A Deformable XAI Music Prototype for Non-Musicians

Ziqing Xu, Nick Bryan-Kinns

机构 * Creative Computing Institute, University of the Arts London(创意计算研究所,伦敦艺术大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments In Proceedings of Explainable AI for the Arts Workshop 2025 (XAIxArts 2025) arXiv:2406.14485

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14915 2025-07-30 cs.MM cs.SD eess.AS 62%

Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling

Xiaojie Li, Ronghui Li, Shukai Fang, Shuzhao Xie, Xiaoyang Guo, Jiaqing Zhou, Junkun Peng, Zhi Wang

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ByteDance Games(字节跳动游戏)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19356 2025-07-28 cs.CL cs.SD eess.AS 62%

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

Hsuan-Yu Wang, Pei-Ying Lee, Berlin Chen

机构 * Department of English(英语系) National Taiwan Normal University(台湾师范大学) Department of Computer Science and Information Engineering(计算机科学与信息工程系)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments 6 pages, 3 figures, to appear in the Proceedings of the 2025 International Conference on Asian Language Processing (IALP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06656 2025-07-22 eess.AS cs.CL cs.LG cs.SD 62%

Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Taejin Park, Ivan Medennikov, Kunal Dhawan, Weiqing Wang, He Huang, Nithin Rao Koluguri, Krishna C. Puvvada, Jagadeesh Balam, Boris Ginsburg

机构 * nvidia

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Published at ICML 2025

Journal ref Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13264 2025-07-18 cs.SD cs.AI eess.AS 62%

Voxtral

Alexander H. Liu, Andy Ehrenberg, Andy Lo, Clément Denoix, Corentin Barreau, Guillaume Lample, Jean-Malo Delignon, Khyathi Raghavi Chandu, Patrick von Platen, Pavankumar Reddy Muddireddy, Sanchit Gandhi, Soham Ghosh, Srijan Mishra, Thomas Foubert, Abhinav Rastogi, Adam Yang, Albert Q. Jiang, Alexandre Sablayrolles, Amélie Héliou, Amélie Martin, Anmol Agarwal, Antoine Roux, Arthur Darcet, Arthur Mensch, Baptiste Bout, Baptiste Rozière, Baudouin De Monicault, Chris Bamford, Christian Wallenwein, Christophe Renaudin, Clémence Lanfranchi, Darius Dabert, Devendra Singh Chaplot, Devon Mizelle, Diego de las Casas, Elliot Chane-Sane, Emilien Fugier, Emma Bou Hanna, Gabrielle Berrada, Gauthier Delerce, Gauthier Guinet, Georgii Novikov, Guillaume Martin, Himanshu Jaju, Jan Ludziejewski, Jason Rute, Jean-Hadrien Chabran, Jessica Chudnovsky, Joachim Studnia, Joep Barmentlo, Jonas Amar, Josselin Somerville Roberts, Julien Denize, Karan Saxena, Karmesh Yadav, Kartik Khandelwal, Kush Jain, Lélio Renard Lavaud, Léonard Blier, Lingxiao Zhao, Louis Martin, Lucile Saulnier, Luyu Gao, Marie Pellat, Mathilde Guillaumin, Mathis Felardos, Matthieu Dinot, Maxime Darrin, Maximilian Augustin, Mickaël Seznec, Neha Gupta, Nikhil Raghuraman, Olivier Duchenne, Patricia Wang, Patryk Saffer, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Philomène Chagniot, Pierre Stock, Pravesh Agrawal, Rémi Delacourt, Romain Sauvestre, Roman Soletskyi, Sagar Vaze, Sandeep Subramanian, Saurabh Garg, Shashwat Dalal, Siddharth Gandhi, Sumukh Aithal, Szymon Antoniak, Teven Le Scao, Thibault Schueller, Thibaut Lavril, Thomas Robert, Thomas Wang, Timothée Lacroix, Tom Bewley, Valeriia Nemychnikova, Victor Paltz, Virgile Richard, Wen-Ding Li, William Marshall, Xuanyu Zhang, Yihan Wan, Yunhao Tang

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00646 2025-07-17 cs.SD cs.AI cs.LG eess.AS 62%

Epic-Sounds: A Large-scale Dataset of Actions That Sound

Jaesung Huh, Jacob Chalk, Evangelos Kazakos, Dima Damen, Andrew Zisserman

机构 * Visual Geometry Group, Department of Engineering Science, University of Oxford, UK(牛津大学视觉几何组) Department of Computer Science, University of Bristol, UK(布里斯托大学计算机科学系) CIIRC, Czech Technical University in Prague, Czech Republic(布拉格捷克技术大学CIIRC)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS

Comments Accepted at TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13766 2025-07-15 cs.SD cs.AI eess.AS 62%

Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge

Ruiyang Qin, Dancheng Liu, Gelei Xu, Zheyu Yan, Chenhui Xu, Yuting Hu, Shaocong Wang, X. Sharon Hu, Jinjun Xiong, Yiyu Shi

机构 * Villanova University(维拉诺瓦大学) University of Notre Dame(诺特尔大学) University at Buffalo–SUNY(布法罗大学–SUNY)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted by ICCAD'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05894 2025-07-09 cs.AI cs.CL 62%

MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation

Fathinah Izzati, Xinyue Li, Yuxuan Wu, Gus Xia

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04182 2025-07-08 cs.IR cs.CL cs.HC cs.SD eess.AS 62%

Navigating Speech Recording Collections with AI-Generated Illustrations

Sirina Håland, Trond Karlsen Strøm, Petra Galuščáková

机构 * University of Stavanger(斯塔万格大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Journal ref SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18055 2025-06-24 cs.MM cs.SD eess.AS 62%

Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings

Jason Clarke, Yoshihiko Gotoh, Stefan Goetze

机构 * Speech and Hearing (SPandH) group, School of Computer Science, The University of Sheffield, Sheffield, United Kingdom(语音与听力(SPandH)小组,计算机科学学院,谢菲尔德大学,谢菲尔德,英国) South Westphalia University of Applied Sciences, Iserlohn, Germany(西南弗劳恩霍夫应用科学大学,伊塞尔洛恩,德国)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM、eess.AS

Comments Accepted to EUSIPCO 2025. 5 pages, 1 figure. To appear in the Proceedings of the 33rd European Signal Processing Conference (EUSIPCO), September 8-12, 2025, Palermo, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16251 2025-06-23 cs.CL eess.AS 62%

End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data

Aishwarya Pothula, Bhavana Akkiraju, Srihari Bandarupalli, Charan D, Santosh Kesiraju, Anil Kumar Vuppala

机构 * Speech Processing Laboratory(语音处理实验室)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20007 2025-06-23 cs.AI cs.CV 62%

Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage

Anita Srbinovska, Angela Srbinovska, Vivek Senthil, Adrian Martin, John McCluskey, Jonathan Bateman, Ernest Fokoué

机构 * Department of Computer Science, Rochester Institute of Technology(罗切斯特理工学院计算机科学系) School of Information, Rochester Institute of Technology(罗切斯特理工学院信息学院) Office of Business Intelligence, Rochester Police Department(罗切斯特警察局商务智能办公室) School of Criminal Justice, University at Albany(阿尔巴尼大学犯罪学学院) School of Individualized Study, Rochester Institute of Technology(罗切斯特理工学院个性化研究学院) Department of Sociology and Anthropology, Rochester Institute of Technology(罗切斯特理工学院社会学与人类学系) School of Mathematics and Statistics, Rochester Institute of Technology(罗切斯特理工学院数学与统计学学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 7 pages, 3 figures, and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13833 2025-06-18 cs.SD cs.AI cs.RO eess.AS physics.app-ph 62%

A Survey on World Models Grounded in Acoustic Physical Information

Xiaoliang Chen, Le Chang, Xin Yu, Yunhe Huang, Xianling Tu

机构 * SoundAI Technology Co., Ltd.(声AI技术有限公司)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 28 pages,11 equations

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12222 2025-06-17 cs.SD cs.AI cs.LG eess.AS 62%

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

Tony Alex, Sara Ahmed, Armin Mustafa, Muhammad Awais, Philip JB Jackson

机构 * Surrey Institute for People-Centred AI(萨里人本人工智能研究所) University of Surrey(萨里大学) Centre for Vision, Speech and Signal Processing (CVSSP)(视觉、语音与信号处理中心)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted at ICLR 2025. Code and pre-trained models are available at \url{https://github.com/ta012/SSLAM}

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12199 2025-06-17 cs.SD cs.AI eess.AS 62%

ViSAGe: Video-to-Spatial Audio Generation

Jaeyeon Kim, Heeseung Yun, Gunhee Kim

机构 * Seoul National University(首尔国立大学)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS

Comments ICLR 2025. Project page: https://jaeyeonkim99.github.io/visage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12008 2025-06-16 cs.SD cs.AI cs.HC eess.AS 62%

Reimagining Dance: Real-time Music Co-creation between Dancers and AI

Olga Vechtomova, Jeff Bos

机构 * University of Waterloo(滑铁卢大学) WordSynth Inc.(WordSynth公司)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted for publication at ICCC 2025 (International Conference on Computational Creativity)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11072 2025-06-16 eess.AS cs.CL cs.CY cs.SD stat.AP 62%

Can We Trust Machine Learning? The Reliability of Features from Open-Source Speech Analysis Tools for Speech Modeling

Tahiya Chowdhury, Veronica Romero

机构 * Department of Computer Science(计算机科学系) Department of Psychology(心理学系)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL、eess.AS

Comments 5 pages, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09709 2025-06-12 cs.SD cs.CV cs.LG eess.AS 62%

Training-Free Voice Conversion with Factorized Optimal Transport

Alexander Lobashev, Assel Yermekova, Maria Larchenko

机构 * Glam AISan Francisco, US Independent researcher(独立研究者) Astana, Kazakhstan Magicly AI Dubai, UAE

专题命中 音频语音多模态 :any-to-any(abstract);分类 cs.CV、eess.AS

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14259 2025-06-12 cs.CL cs.AI 62%

Let's Fuse Step by Step: A Generative Fusion Decoding Algorithm with LLMs for Robust and Instruction-Aware ASR and OCR

Chan-Jan Hsu, Yi-Chang Chen, Feng-Ting Liao, Pei-Chen Ho, Yu-Hsiang Wang, Po-Chun Hsu, Da-shan Shiu

机构 * MediaTek Research(联发科研究)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏