发表机构
LTCI, Télécom Paris, Institut Polytechnique de Paris(巴黎理工学院电信巴黎高等学院LTCI实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
我们提出了互联网档案馆音乐数据集(IAMD),一个包含超过34,000小时音频的最大公开音乐-字幕数据集,并设计了一种自动字幕生成流程,经评估该流程可靠且可扩展。
AI 中文摘要
我们推出了互联网档案馆音乐数据集(IAMD),这是一个从互联网档案馆获取的大规模带字幕音乐片段集合。据我们所知,IAMD 是迄今为止最大的公开可用的音乐-字幕数据集,包含超过 34,000 小时的音频,为训练和评估音乐理解与生成模型提供了宝贵的基准。该数据集基于声明以知识共享许可协议分发的内容构建,并与 MusicBrainz 进行交叉引用,以提高许可信息的可靠性。为了对 IAMD 进行标注,我们提出了一种自动字幕生成流程,该流程通过音频语言模型(ALM)生成基础字幕,并利用来自互联网档案馆的文本元数据以及通过音频分类模型获得的插补元数据对其进行增强。字幕质量通过客观和主观方式评估,结果表明该标注流程可靠,且随着规模扩大不会降低字幕质量。
英文摘要
We introduce the Internet Archive Music Dataset (IAMD), a large-scale collection of captioned music segments derived from the Internet Archive. To the best of our knowledge, IAMD constitutes the largest publicly available music-caption dataset to date with over 34,000 hours of audio, providing a valuable benchmark for training and evaluating music understanding and generative models. The dataset is built from content declared to be distributed under Creative Commons licenses, and cross-referencing with MusicBrainz is done to improve license information reliability. To annotate IAMD, we present an automatic captioning pipeline that augments base captions produced by an audio-language model (ALM) with textual metadata sourced from the Internet Archive and imputed metadata obtained using audio classification models. Caption quality is assessed objectively and subjectively, and results indicate that the annotation pipeline is reliable and does not degrade caption quality with scaling.
Journal refISMIR, 2026, ABU DHABI, United Arab Emirates