文本提示的CLAP:通过对比学习学习查询条件音频表示
Text-Prompted CLAP: Learning Text-Conditioned Audio Representations via Contrastive Learning
浏览论文内容
中文总结 AI 辅助
研究旨在解决CLAP独立编码模态限制,提出TP-CLAP,通过引入交叉注意力融合模块,利用音频多项选择题框架训练,在音频问答、检索等任务中表现出色,优于标准CLAP基线。
中文摘要 AI 辅助
对比语言-音频预训练(CLAP)在共享嵌入空间中学习对齐的文本和音频表示。然而,每个模态的独立编码限制了其在复杂音频理解和检索任务中对跨模态语义建模的能力。为解决此限制,本文提出文本提示的CLAP(TP-CLAP),它是CLAP的参数高效扩展,引入基于交叉注意力的融合模块将文本提示纳入音频特征。TP-CLAP使用音频多项选择题框架训练,通过对比学习使查询条件音频表示与正确答案选择的文本嵌入对齐。实验表明,TP-CLAP在音频问答(audio-QA)中使用大得多的音频语言模型取得了有竞争力的性能,同时在音频-文本检索和零样本分类基准上改进了基础CLAP模型。学习到的表示进一步针对属性聚焦的音频到音频检索进行微调,表明TP-CLAP在音乐检索任务中始终优于标准CLAP基线。
英文摘要
Contrastive Language-Audio Pretraining (CLAP) aligns text and audio in a shared embedding space, but encoding each modality independently limits its ability to model cross-modal semantics in complex audio understanding and retrieval tasks. To address this limitation, this paper proposes Text-Prompted CLAP (TP-CLAP), a parameter-efficient extension of CLAP that introduces a cross-attention-based fusion module to incorporate textual prompts into audio features. TP-CLAP is trained using an audio multiple-choice question answering (AMCQA) framework, where it learns to align text-conditioned audio representations with text embeddings of correct answer choices via contrastive learning. Experiments demonstrate that TP-CLAP matches or exceeds several substantially larger audio-LLMs on audio question answering despite its compact size. After fine-tuning for attribute-focused audio-to-audio retrieval, text-conditioned audio representations consistently outperform their unconditioned counterparts on music retrieval benchmarks. In addition, TP-CLAP improves upon the base CLAP model on conventional audio-text retrieval and zero-shot classification.