SonicCaps:用于改进音频检索的大规模多样化细粒度字幕数据集
SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval
浏览论文内容
中文总结 AI 辅助
该研究提出大规模音频字幕数据集SonicCaps,通过多模态大语言模型Qwen3-Omni生成约1500万条字幕,采用多字幕采样策略训练CLAP模型,提升了音频检索等任务的性能与泛化能力。
中文摘要 AI 辅助
音频语言建模的近期进展得益于大规模音频字幕数据集的推动。然而,现有数据集仍受限于语义多样性低、缺乏声学细节的通用描述,以及一对一的音频-字幕映射,这些映射无法很好地反映听觉感知固有的模糊性。我们引入SonicCaps,这是一个包含约1500万条字幕和约70万个音频片段的大规模音频字幕数据集,通过多模态大语言模型Qwen3-Omni在音频和文本的条件下生成。为了明确促进多样性,我们通过结构化提示工程和少样本生成,为每个音频生成约24条字幕,涵盖主要描述、改写变体(不同详尽程度、风格)和语义标签。人工评估显示,SonicCaps的评分显著高于现有字幕数据集,细粒度分析表明,我们的字幕被认为更具描述性和精确性,这与质量判断密切相关。最后,在SonicCaps上采用多字幕采样策略训练CLAP模型,可持续改进音频检索和零样本分类,在公共和商业基准上具有更强的泛化能力。我们在hugging face上发布SonicCaps和两个专门的CLAP模型:this https URL
英文摘要
Recent advances in audio-language modeling have been driven by large-scale audio captioning datasets. However, existing datasets remain limited by low semantic diversity, generic descriptions lacking acoustic details, and one-to-one audio-caption mappings that poorly reflect the inherent ambiguity of auditory perception. We introduce SonicCaps, a large-scale audio captioning dataset comprising ~15M captions paired with ~700k audio clips, generated using a multi-modal large language model (Qwen3-Omni) conditioned on both audio and text. To explicitly promote diversity, we generate around 24 captions per audio via structured prompt engineering and few- shot generation, spanning main descriptions, rephrased variants (verbosity, style) and semantic tags. Human evaluation shows that SonicCaps is rated significantly higher than existing captioning datasets, with fine-grained analyses indicating that our captions are perceived as more descriptive and precise, which strongly correlates with quality judgments. Finally, training CLAP models on SonicCaps with a multi-caption sampling strategy consistently improves audio retrieval and zero-shot classification, with stronger generalization across public and commercial benchmarks. We release both SonicCaps and two specialized CLAP models on hugging face: https://huggingface.co/datasets/Zineb/SonicCaps.
发表机构
- Sony CTC(索尼CTC)
- LTCI, Telecom Paris, Institut polytechnique de Paris(巴黎综合理工学院电信学院LTCI)
机构由 AI 辅助整理,请以论文原文为准。