SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
机构 * SimulaMet ; OsloMet ; Forzasys ; University of Central Florida(佛罗里达中央大学) ; University of Liège(列日大学) ; KAUST(科威特大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * SimulaMet ; OsloMet ; Forzasys ; University of Central Florida(佛罗里达中央大学) ; University of Liège(列日大学) ; KAUST(科威特大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS
机构 * University of Southern California(南加州大学) ; Stanford University(斯坦福大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS
机构 * China Mobile Internet Company Ltd.(中国移动互联网有限公司) ; Northeastern University(东北大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.MM
机构 * Massachusetts Institute of Technology(麻省理工学院)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS
Comments Accepted to ASRU 2025
机构 * Center For Artificial Intelligence and Data Science, University of Würzburg(人工智能与数据科学中心,乌尔姆大学) ; Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学) ; Mila, McGill University and Canada CIFAR AI Chair(Mila,麦吉尔大学及加拿大CIFAR人工智能主席)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS
Comments Accepted to the 26th International Society for Music Information Retrieval Conference (ISMIR), 2025
机构 * Department of MetaBioHealth, Sungkyunkwan University, Korea(韩国成均馆大学代谢生物健康系) ; Department of Applied Artificial Intelligence, Sungkyunkwan University, Korea(韩国成均馆大学应用人工智能系) ; Department of Computer Science and Engineering, Jaume I University, Spain(西班牙伊萨贝拉大学计算机科学与工程系)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、cs.MM
Comments Accepted for publication at IJCAI 2025. 9 pages, 4 tables, 3 figures
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; Tencent AI Lab(腾讯人工智能实验室)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.MM、eess.AS
机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) ; Google(谷歌)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV、eess.AS
Comments 15 pages, 6 figres, 6 tables. Accepted to ISMAR 2025 as a TVCG journal paper
机构 * Johns Hopkins University(约翰霍普金斯大学) ; Southeast University(东南大学) ; National University of Defense Technology(国防科技大学) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、cs.MM
Comments WASPAA 2025
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Tencent AI Lab(腾讯AI实验室)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、eess.AS
Comments 9 pages
机构 * Creative Computing Institute, University of the Arts London(创意计算研究所,伦敦艺术大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS
Comments In Proceedings of Explainable AI for the Arts Workshop 2025 (XAIxArts 2025) arXiv:2406.14485
机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; ByteDance Games(字节跳动游戏)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM、eess.AS
机构 * Department of English(英语系) ; National Taiwan Normal University(台湾师范大学) ; Department of Computer Science and Information Engineering(计算机科学与信息工程系)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS
Comments 6 pages, 3 figures, to appear in the Proceedings of the 2025 International Conference on Asian Language Processing (IALP)
机构 * nvidia
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS
Comments Published at ICML 2025
Journal ref Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS
Comments 17 pages
机构 * Visual Geometry Group, Department of Engineering Science, University of Oxford, UK(牛津大学视觉几何组) ; Department of Computer Science, University of Bristol, UK(布里斯托大学计算机科学系) ; CIIRC, Czech Technical University in Prague, Czech Republic(布拉格捷克技术大学CIIRC)
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS
Comments Accepted at TPAMI
机构 * Villanova University(维拉诺瓦大学) ; University of Notre Dame(诺特尔大学) ; University at Buffalo–SUNY(布法罗大学–SUNY)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI、eess.AS
Comments Accepted by ICCAD'25
机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI
机构 * University of Stavanger(斯塔万格大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS
Journal ref SIGIR 2025
机构 * Speech and Hearing (SPandH) group, School of Computer Science, The University of Sheffield, Sheffield, United Kingdom(语音与听力(SPandH)小组,计算机科学学院,谢菲尔德大学,谢菲尔德,英国) ; South Westphalia University of Applied Sciences, Iserlohn, Germany(西南弗劳恩霍夫应用科学大学,伊塞尔洛恩,德国)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM、eess.AS
Comments Accepted to EUSIPCO 2025. 5 pages, 1 figure. To appear in the Proceedings of the 33rd European Signal Processing Conference (EUSIPCO), September 8-12, 2025, Palermo, Italy
机构 * Speech Processing Laboratory(语音处理实验室)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL、eess.AS
机构 * Department of Computer Science, Rochester Institute of Technology(罗切斯特理工学院计算机科学系) ; School of Information, Rochester Institute of Technology(罗切斯特理工学院信息学院) ; Office of Business Intelligence, Rochester Police Department(罗切斯特警察局商务智能办公室) ; School of Criminal Justice, University at Albany(阿尔巴尼大学犯罪学学院) ; School of Individualized Study, Rochester Institute of Technology(罗切斯特理工学院个性化研究学院) ; Department of Sociology and Anthropology, Rochester Institute of Technology(罗切斯特理工学院社会学与人类学系) ; School of Mathematics and Statistics, Rochester Institute of Technology(罗切斯特理工学院数学与统计学学院)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 7 pages, 3 figures, and 1 table
机构 * SoundAI Technology Co., Ltd.(声AI技术有限公司)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS
Comments 28 pages,11 equations
机构 * Surrey Institute for People-Centred AI(萨里人本人工智能研究所) ; University of Surrey(萨里大学) ; Centre for Vision, Speech and Signal Processing (CVSSP)(视觉、语音与信号处理中心)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS
Comments Accepted at ICLR 2025. Code and pre-trained models are available at \url{https://github.com/ta012/SSLAM}
机构 * Seoul National University(首尔国立大学)
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.AI、eess.AS
Comments ICLR 2025. Project page: https://jaeyeonkim99.github.io/visage/
机构 * University of Waterloo(滑铁卢大学) ; WordSynth Inc.(WordSynth公司)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI、eess.AS
Comments Accepted for publication at ICCC 2025 (International Conference on Computational Creativity)
机构 * Department of Computer Science(计算机科学系) ; Department of Psychology(心理学系)
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CL、eess.AS
Comments 5 pages, 1 figure, 3 tables
机构 * Glam AISan Francisco, US ; Independent researcher(独立研究者) ; Astana, Kazakhstan ; Magicly AI ; Dubai, UAE
专题命中 音频语音多模态 :any-to-any(abstract);分类 cs.CV、eess.AS
Comments Interspeech 2025
机构 * MediaTek Research(联发科研究)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI