arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4597 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4597 篇

2508.01915 2025-08-05 cs.CV cs.ET cs.HC cs.LG cs.SD eess.AS 62%

EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses

Akshay Paruchuri, Sinan Hersek, Lavisha Aggarwal, Qiao Yang, Xin Liu, Achin Kulshrestha, Andrea Colaco, Henry Fuchs, Ishan Chatterjee

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Google(谷歌)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、eess.AS

Comments 15 pages, 6 figres, 6 tables. Accepted to ISMAR 2025 as a TVCG journal paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06690 2025-08-05 cs.SD cs.AI cs.MM 62%

Benchmarking Sub-Genre Classification For Mainstage Dance Music

Hongzhi Shu, Xinglin Li, Hongyu Jiang, Minghao Fu, Xinyu Li

机构 * Johns Hopkins University(约翰霍普金斯大学) Southeast University(东南大学) National University of Defense Technology(国防科技大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、cs.MM

Comments WASPAA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00391 2025-08-04 cs.CV eess.AS 62%

Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition

Guanjie Huang, Danny H. K. Tsang, Shan Yang, Guangzhi Lei, Li Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tencent AI Lab(腾讯AI实验室)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、eess.AS

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00160 2025-08-04 cs.HC cs.AI cs.SD eess.AS 62%

DeformTune: A Deformable XAI Music Prototype for Non-Musicians

Ziqing Xu, Nick Bryan-Kinns

机构 * Creative Computing Institute, University of the Arts London(创意计算研究所,伦敦艺术大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments In Proceedings of Explainable AI for the Arts Workshop 2025 (XAIxArts 2025) arXiv:2406.14485

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14915 2025-07-30 cs.MM cs.SD eess.AS 62%

Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling

Xiaojie Li, Ronghui Li, Shukai Fang, Shuzhao Xie, Xiaoyang Guo, Jiaqing Zhou, Junkun Peng, Zhi Wang

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ByteDance Games(字节跳动游戏)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19356 2025-07-28 cs.CL cs.SD eess.AS 62%

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

Hsuan-Yu Wang, Pei-Ying Lee, Berlin Chen

机构 * Department of English(英语系) National Taiwan Normal University(台湾师范大学) Department of Computer Science and Information Engineering(计算机科学与信息工程系)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments 6 pages, 3 figures, to appear in the Proceedings of the 2025 International Conference on Asian Language Processing (IALP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.06656 2025-07-22 eess.AS cs.CL cs.LG cs.SD 62%

Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

Taejin Park, Ivan Medennikov, Kunal Dhawan, Weiqing Wang, He Huang, Nithin Rao Koluguri, Krishna C. Puvvada, Jagadeesh Balam, Boris Ginsburg

机构 * nvidia

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Published at ICML 2025

Journal ref Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13264 2025-07-18 cs.SD cs.AI eess.AS 62%

Voxtral

Alexander H. Liu, Andy Ehrenberg, Andy Lo, Clément Denoix, Corentin Barreau, Guillaume Lample, Jean-Malo Delignon, Khyathi Raghavi Chandu, Patrick von Platen, Pavankumar Reddy Muddireddy, Sanchit Gandhi, Soham Ghosh, Srijan Mishra, Thomas Foubert, Abhinav Rastogi, Adam Yang, Albert Q. Jiang, Alexandre Sablayrolles, Amélie Héliou, Amélie Martin, Anmol Agarwal, Antoine Roux, Arthur Darcet, Arthur Mensch, Baptiste Bout, Baptiste Rozière, Baudouin De Monicault, Chris Bamford, Christian Wallenwein, Christophe Renaudin, Clémence Lanfranchi, Darius Dabert, Devendra Singh Chaplot, Devon Mizelle, Diego de las Casas, Elliot Chane-Sane, Emilien Fugier, Emma Bou Hanna, Gabrielle Berrada, Gauthier Delerce, Gauthier Guinet, Georgii Novikov, Guillaume Martin, Himanshu Jaju, Jan Ludziejewski, Jason Rute, Jean-Hadrien Chabran, Jessica Chudnovsky, Joachim Studnia, Joep Barmentlo, Jonas Amar, Josselin Somerville Roberts, Julien Denize, Karan Saxena, Karmesh Yadav, Kartik Khandelwal, Kush Jain, Lélio Renard Lavaud, Léonard Blier, Lingxiao Zhao, Louis Martin, Lucile Saulnier, Luyu Gao, Marie Pellat, Mathilde Guillaumin, Mathis Felardos, Matthieu Dinot, Maxime Darrin, Maximilian Augustin, Mickaël Seznec, Neha Gupta, Nikhil Raghuraman, Olivier Duchenne, Patricia Wang, Patryk Saffer, Paul Jacob, Paul Wambergue, Paula Kurylowicz, Philomène Chagniot, Pierre Stock, Pravesh Agrawal, Rémi Delacourt, Romain Sauvestre, Roman Soletskyi, Sagar Vaze, Sandeep Subramanian, Saurabh Garg, Shashwat Dalal, Siddharth Gandhi, Sumukh Aithal, Szymon Antoniak, Teven Le Scao, Thibault Schueller, Thibaut Lavril, Thomas Robert, Thomas Wang, Timothée Lacroix, Tom Bewley, Valeriia Nemychnikova, Victor Paltz, Virgile Richard, Wen-Ding Li, William Marshall, Xuanyu Zhang, Yihan Wan, Yunhao Tang

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 17 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00646 2025-07-17 cs.SD cs.AI cs.LG eess.AS 62%

Epic-Sounds: A Large-scale Dataset of Actions That Sound

Jaesung Huh, Jacob Chalk, Evangelos Kazakos, Dima Damen, Andrew Zisserman

机构 * Visual Geometry Group, Department of Engineering Science, University of Oxford, UK(牛津大学视觉几何组) Department of Computer Science, University of Bristol, UK(布里斯托大学计算机科学系) CIIRC, Czech Technical University in Prague, Czech Republic(布拉格捷克技术大学CIIRC)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS

Comments Accepted at TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13766 2025-07-15 cs.SD cs.AI eess.AS 62%

Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge

Ruiyang Qin, Dancheng Liu, Gelei Xu, Zheyu Yan, Chenhui Xu, Yuting Hu, Shaocong Wang, X. Sharon Hu, Jinjun Xiong, Yiyu Shi

机构 * Villanova University(维拉诺瓦大学) University of Notre Dame(诺特尔大学) University at Buffalo–SUNY(布法罗大学–SUNY)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted by ICCAD'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05894 2025-07-09 cs.AI cs.CL 62%

MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation

Fathinah Izzati, Xinyue Li, Yuxuan Wu, Gus Xia

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04182 2025-07-08 cs.IR cs.CL cs.HC cs.SD eess.AS 62%

Navigating Speech Recording Collections with AI-Generated Illustrations

Sirina Håland, Trond Karlsen Strøm, Petra Galuščáková

机构 * University of Stavanger(斯塔万格大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Journal ref SIGIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18055 2025-06-24 cs.MM cs.SD eess.AS 62%

Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings

Jason Clarke, Yoshihiko Gotoh, Stefan Goetze

机构 * Speech and Hearing (SPandH) group, School of Computer Science, The University of Sheffield, Sheffield, United Kingdom(语音与听力(SPandH)小组,计算机科学学院,谢菲尔德大学,谢菲尔德,英国) South Westphalia University of Applied Sciences, Iserlohn, Germany(西南弗劳恩霍夫应用科学大学,伊塞尔洛恩,德国)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM、eess.AS

Comments Accepted to EUSIPCO 2025. 5 pages, 1 figure. To appear in the Proceedings of the 33rd European Signal Processing Conference (EUSIPCO), September 8-12, 2025, Palermo, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16251 2025-06-23 cs.CL eess.AS 62%

End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data

Aishwarya Pothula, Bhavana Akkiraju, Srihari Bandarupalli, Charan D, Santosh Kesiraju, Anil Kumar Vuppala

机构 * Speech Processing Laboratory(语音处理实验室)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20007 2025-06-23 cs.AI cs.CV 62%

Towards AI-Driven Policing: Interdisciplinary Knowledge Discovery from Police Body-Worn Camera Footage

Anita Srbinovska, Angela Srbinovska, Vivek Senthil, Adrian Martin, John McCluskey, Jonathan Bateman, Ernest Fokoué

机构 * Department of Computer Science, Rochester Institute of Technology(罗切斯特理工学院计算机科学系) School of Information, Rochester Institute of Technology(罗切斯特理工学院信息学院) Office of Business Intelligence, Rochester Police Department(罗切斯特警察局商务智能办公室) School of Criminal Justice, University at Albany(阿尔巴尼大学犯罪学学院) School of Individualized Study, Rochester Institute of Technology(罗切斯特理工学院个性化研究学院) Department of Sociology and Anthropology, Rochester Institute of Technology(罗切斯特理工学院社会学与人类学系) School of Mathematics and Statistics, Rochester Institute of Technology(罗切斯特理工学院数学与统计学学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 7 pages, 3 figures, and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13833 2025-06-18 cs.SD cs.AI cs.RO eess.AS physics.app-ph 62%

A Survey on World Models Grounded in Acoustic Physical Information

Xiaoliang Chen, Le Chang, Xin Yu, Yunhe Huang, Xianling Tu

机构 * SoundAI Technology Co., Ltd.(声AI技术有限公司)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

Comments 28 pages,11 equations

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12222 2025-06-17 cs.SD cs.AI cs.LG eess.AS 62%

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

Tony Alex, Sara Ahmed, Armin Mustafa, Muhammad Awais, Philip JB Jackson

机构 * Surrey Institute for People-Centred AI(萨里人本人工智能研究所) University of Surrey(萨里大学) Centre for Vision, Speech and Signal Processing (CVSSP)(视觉、语音与信号处理中心)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted at ICLR 2025. Code and pre-trained models are available at \url{https://github.com/ta012/SSLAM}

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12199 2025-06-17 cs.SD cs.AI eess.AS 62%

ViSAGe: Video-to-Spatial Audio Generation

Jaeyeon Kim, Heeseung Yun, Gunhee Kim

机构 * Seoul National University(首尔国立大学)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS

Comments ICLR 2025. Project page: https://jaeyeonkim99.github.io/visage/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12008 2025-06-16 cs.SD cs.AI cs.HC eess.AS 62%

Reimagining Dance: Real-time Music Co-creation between Dancers and AI

Olga Vechtomova, Jeff Bos

机构 * University of Waterloo(滑铁卢大学) WordSynth Inc.(WordSynth公司)

专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS

Comments Accepted for publication at ICCC 2025 (International Conference on Computational Creativity)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11072 2025-06-16 eess.AS cs.CL cs.CY cs.SD stat.AP 62%

Can We Trust Machine Learning? The Reliability of Features from Open-Source Speech Analysis Tools for Speech Modeling

Tahiya Chowdhury, Veronica Romero

机构 * Department of Computer Science(计算机科学系) Department of Psychology(心理学系)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL、eess.AS

Comments 5 pages, 1 figure, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09709 2025-06-12 cs.SD cs.CV cs.LG eess.AS 62%

Training-Free Voice Conversion with Factorized Optimal Transport

Alexander Lobashev, Assel Yermekova, Maria Larchenko

机构 * Glam AISan Francisco, US Independent researcher(独立研究者) Astana, Kazakhstan Magicly AI Dubai, UAE

专题命中 音频语音多模态 :any-to-any(abstract);分类 cs.CV、eess.AS

Comments Interspeech 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.14259 2025-06-12 cs.CL cs.AI 62%

Let's Fuse Step by Step: A Generative Fusion Decoding Algorithm with LLMs for Robust and Instruction-Aware ASR and OCR

Chan-Jan Hsu, Yi-Chang Chen, Feng-Ting Liao, Pei-Chen Ho, Yu-Hsiang Wang, Po-Chun Hsu, Da-shan Shiu

机构 * MediaTek Research(联发科研究)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03710 2025-06-12 eess.AS cs.CL cs.SD 62%

Listen, Chat, and Remix: Text-Guided Soundscape Remixing for Enhanced Auditory Experience

Xilin Jiang, Cong Han, Yinghao Aaron Li, Nima Mesgarani

机构 * Department of Electrical Engineering, Columbia University(电气工程系,哥伦比亚大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted by IEEE Journal of Selected Topics in Signal Processing (JSTSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08634 2025-06-11 cs.HC cs.AI cs.CV 62%

MOSAIC-F: A Framework for Enhancing Students' Oral Presentation Skills through Personalized Feedback

Alvaro Becerra, Daniel Andres, Pablo Villegas, Roberto Daza, Ruth Cobos

机构 * GHIA Group, School of Engineering, Universidad Autónoma de Madrid, Spain(GHIA小组,工程学院,马德里自治大学,西班牙) BiDA-Lab Group, School of Engineering, Universidad Autónoma de Madrid, Spain(BiDA实验室小组,工程学院,马德里自治大学,西班牙)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted in LASI Spain 25: Learning Analytics Summer Institute Spain 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08279 2025-06-11 cs.CV cs.AI cs.LG 62%

Seeing Voices: Generating A-Roll Video from Audio with Mirage

Aditi Sundararaman, Amogh Adishesha, Andrew Jaegle, Dan Bigioi, Hyoung-Kyu Song, Jon Kyl, Justin Mao, Kevin Lan, Mojtaba Komeili, ShahRukh Athar, Sheila Babayan, Stanislau Beliasau, William Buchwalter

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Technical report website: mirage.app/research/seeing-voices, product website: mirage.app

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07572 2025-06-10 cs.CV cs.CL 62%

Learning Speaker-Invariant Visual Features for Lipreading

Yu Li, Feng Xue, Shujie Li, Jinrui Zhang, Shuang Yang, Dan Guo, Richang Hong

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06537 2025-06-10 cs.CV cs.SD eess.AS 62%

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

Seung-jae Lee, Paul Hongsuck Seo

机构 * Department of Computer Science and Engineering(计算机科学与工程系)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、eess.AS

Comments Accepted on INTERSPEECH2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03214 2025-06-05 q-bio.NC cs.AI cs.CL 62%

A Pre-trained Framework for Multilingual Brain Decoding Using Non-invasive Recordings

Yi Guo, Yihang Dong, Michael Kwok-Po Ng, Shuqiang Wang

机构 * Shenzhen Institute of Advanced Technology(深圳先进技术研究院) Chinese Academy of Sciences(中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Hong Kong Baptist University(香港 Baptist大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.11607 2025-06-05 cs.CL cs.SD eess.AS 62%

Transformers in Speech Processing: A Survey

Siddique Latif, Aun Zaidi, Heriberto Cuayahuitl, Fahad Shamshad, Moazzam Shoukat, Muhammad Usama, Junaid Qadir

机构 * Queensland University of Technology(昆士兰理工大学) Information Technology University(信息科技大学) University of Lincoln(林肯大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) National University of Computer and Emerging Sciences(国家计算机与新兴科学大学) Qatar University(卡塔尔大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted in Computer Science Review 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00722 2025-06-03 cs.CL cs.SD eess.AS 62%

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

Siddhant Arora, Jinchuan Tian, Hayato Futami, Jee-weon Jung, Jiatong Shi, Yosuke Kashiwagi, Emiru Tsunoo, Shinji Watanabe

机构 * Carnegie Mellon University(卡内基梅隆大学) Sony Group Corporation(索尼集团)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS

Comments Accepted at INTERSPEECH 2025

详情

展开后加载摘要…

URL PDF HTML 收藏